Generating and visualising multivariate random numbers in R

This post will present the wonderful pairs.panels function of the psych package [1] that I discovered recently to visualise multivariate random numbers.

Here is a little example with a Gaussian copula and normal and log-normal marginal distributions. I use pairs.panels to illustrate the steps along the way.

I start with standardised multivariate normal random numbers:
library(psych)
library(MASS)
Sig <- matrix(c(1, -0.7, -.5,
                -0.7, 1, 0.6,
                -0.5, 0.6, 1), 
                nrow=3)
X <- mvrnorm(1000, mu=rep(0,3), Sigma = Sig, 
             empirical = TRUE)  
pairs.panels(X)

Who will win the World Cup and which prediction model?

The World Cup has finally kicked off last Thursday and I have seen some fantastic games already. Perhaps the Netherlands appears to be the strongest side so far, following their 5-1 victory over Spain.

To me the question is not only which country will win the World Cup, but also which prediction model will come closest to the actual results. Here I present three teams, FiveThirtyEight, a polling aggregation website, Groll & Schauberger, two academics from Munich and finally Lloyd's of London, the insurance market.

The guys around Nate Silver at FiveThirtyEight have used the ESPN’s Soccer Power Index to predict the outcomes of games and continue to update their model. Brazil is their clear favourite.

Andreas Groll & Gunther Schauberger from the LMU Munich developed a model, which like the approach from FiveThirtyEight aims to estimate the probability of a team winning the world cup. But unlike FiveThirtyEight, they see Germany to take the trophy home.

Lloyd's chose a different approach for predicting the World Cup final. The insurance market looked at the risk aspect of the teams and ranked the teams by their insured value. Arguably the better a team the higher their insured value. As a result Lloyd's predicts Germany to win the World Cup.

Quick reminder; what's the difference between insurance and gambling? Gambling introduces risk, where none exists. Insurance mitigates risk, where risk exists.

The joy of joining data.tables

The example I present here is a little silly, yet it illustrates how to join tables with data.table in R.

Mapping old data to new data

Categories in general are never fixed, they always change at some point. And then the trouble starts with the data. For example not that long ago we didn't distinguish between smartphones and dumbphones, or video on demand and video rental shops.

I would like to back track price change data for smartphones and online movie rental shops, assuming that their earlier development can be set to the categories they were formerly part of, namely mobile and video rental shops to create indices.

Here is my toy data:



I'd like to create price indices for all products and where data for the new product categories is missing, use the price changes of the old product category.


The data.table package helps here. I start with my original data, convert it into a data.table and create mapping tables. That allows me to add the old product with its price change to the new product. The trick here is to set certain columns as key. Two data tables with the same key can easily be joined. Once I have the price changes for products and years the price index can be added. The plot illustrates the result, the R code is below.

Early bird registration for R in Insurance closes tomorrow

The early bird registration offer for the 2nd R in Insurance conference, 14 July 2014, at Cass Business School closes tomorrow.


This one-day conference will focus once more on applications in insurance and actuarial science that use R, the lingua franca for statistical computation. Topics covered include reserving, pricing, loss modelling, the use of R in a production environment, and more. All topics are to be discussed within the context of using R as a primary tool for insurance risk management, analysis and modelling.

The conference programme consists of invited talks and contributed presentations discussing the wide range of fields in which R is used in insurance:


Attendance of the whole conference is the equivalent of 6.5 hours of CPD for members of the Actuarial Profession.

You can register directly on the conference web site www.rininsurance.com.

The organisers gratefully acknowledge the sponsorship of Mango Solutions, CYBAEA, RStudio and PwC without whom the event wouldn't be possible.

Notes from the Kölner R meeting, 23 May 2014

The 10th Kölner R user meeting took place last Friday at the Institute of Sociology and to celebrate the anniversary we invited Andrie de Vries to join us from Revolution Analytics. Andrie is well known in the R community; he is the co-author of the R for Dummies book and an active contributor on stackoverflow.

Taking R to the Enterprise


Andrie de Vries: Taking R to the Enterprise. Photo: Günter Faes

Andrie talked about how R is finding its way into the enterprise. He argued that R has become the go-to tool for data and statistical analysis in many companies. He observed that over the years the commercial user base evolved from a small expert community with a love for the command line and cutting edge technology into a much wider audience. Although R is still used by the technical experts, who demand more and more power, out-of-memory computations on bigger data, etc. the number of casual R users is growing rapidly. Often the casual R user relies on code written by others as part of a bigger workflow and hence prefers a more standardised user interface.

Andrie illustrated how Revolution Analytics aims to satisfy both user groups. He started with the expert R users by demonstrating Revolution Analytics' RevoScaleR package that allows them to carry out analysis on bigger data. His example followed closely the airline analysis of Joseph Rickert's white paper. Andrie then showed how more complex R code and functions can be integrated (and to some extend hidden) into Alteryx and Tableau via Rserve.

googleVis overview & developments


 googleVis Update
Slides and code available on GitHub

I gave a brief introduction to googleVis and presented recent developments. The googleVis package provides an interface between R and the Google Chart Tools API. It allows users to create web pages with interactive charts based on R data frames and to display them either via the local R HTTP help server or within their own sites, without uploading the data to Google. The best known example is perhaps the motion chart of fertility and life expectancy data from the World Bank as first presented by Hans Rosling. Another popular example is the chart that illustrates the performance of the famous Lloyd's insurance market.

Since version 0.5.0 of googleVis many new chart types have been added to googleVis, including annotation, sankey, calendar and timeline charts. Additionally new markdown vignettes have been included to present the examples also on CRAN. The slides and rmarkdown code are available via GitHub.

Kölsch & Schnitzel

No Kölner R meeting would be complete without a few Kölsch and Schnitzel at the Lux. We were lucky with the weather so that we could sit outside to enjoy drinks, food and networking opportunities.

Kölsch and Schnitzel at the Lux. Photo: Günter Faes

Next Kölner R meeting

The next meeting is scheduled for 12 September 2014.

Please get in touch if you would like to present and share your experience, or indeed if you have a request for a topic you would like to hear more about. For more details see also our Meetup page.

Thanks again to Bernd Weiß for hosting the event and Revolution Analytics for their sponsorship.

Next Kölner R User Meeting: Friday, 23 May 2014

The next Cologne R user group meeting is scheduled for this Friday, 23 May 2014.

To celebrate our 10th meeting we welcome:
Followed by drinks and schnitzel at the Lux.

Further details available on our KölnRUG Meetup site. Please sign up if you would like to come along. Notes from past meetings are available here.


The organisers, Bernd Weiß and Markus Gesmann, gratefully acknowledge the sponsorship of Revolution Analytics, who support the Cologne R user group as part of their vector programme.


View Larger Map

The Wiener takes it all? A review of the 2014 Eurovision results


Saturday's Eurovision Song Contest (ESC) from Copenhagen was hilarious as usual with acts from all over Europe and some more or less sensible gimmicks: a circular piano, a giant hamster wheel, a sea-saw, or indeed a beard and fancy dress.

The results of the ESC were only a little different to what the bookmakers in the UK had predicted before the event started. Sweden was seen as the favourite, followed by Austria, Netherlands, Armenia and the UK.

In the end Conchita Wurst from Austria won with 290 points in front of the acts from the Netherlands with 238, Sweden with 218 and Armenia with 174 points. The UK ended on the 17th rank with only 40 points. Perhaps the reason the English bookies had put the UK in fifth place reflected the bias of their clients towards their home country.

The points were given as a combination of a public tele vote and a jury in each country. But how big was the resemblance between the votes of the juries and the public? Is there much variance between countries?

The detailed voting data are available from the Eurovision site. Out of the 37 countries that participated in the voting, two countries, San Marino and Albania, had no tele voting and one country, Georgia, had no jury. In the remaining 34 countries Austria and the Netherlands made into the top 2 of the jury and public voting results if they were treated independently. Sweden made it into the top 5 of both as well, but the other countries differed.


So, how consistent was the voting of the jury and public for the top 5 in each country?

Well, in only 11 countries out of 34, the jury agreed with the public on more than 2 of the top 5 songs. In 25 countries the jury and public agreed on at least two candidates.

In some cases the differences between public and jury were so wide, that although a candidate was voted into the top 5 of the tele rankings the act wouldn't get any points. The top public winner in Belgium (Armenia), Ireland (Poland), Montenegro (Russia) and the UK (Poland) didn't get any points at all.

Still, getting the top favourites right seems much easier then any of the followers. Or in other words, it is really difficult to produce a hit, a song/act on which many can agree. But when you hear one, it is much easier to identify it as such, something you really like and believe others would like as well.

R Code