KnowledgeHub
Questions
Tags
Users
Search
Alex Rivera
|
Logout
Edit Question
Title
Body
This is a bit of a philosophical question about data.table join syntax. I am finding more and more uses for data.tables, but still learning... The join format X[Y] for data.tables is very concise, handy and efficient, but as far as I can tell, it only supports inner joins and right outer joins. To get a left or full outer join, I need to use merge : X[Y, nomatch = NA] -- all rows in Y -- right outer join (default) X[Y, nomatch = 0] -- only rows with matches in both X and Y -- inner join merge(X, Y, all = TRUE) -- all rows from both X and Y -- full outer join merge(X, Y, all.x = TRUE) -- all rows in X -- left outer join It seems to me that it would be handy if the X[Y] join format supported all 4 types of joins. Is there a reason only two types of joins are supported? For me, the nomatch = 0 and nomatch = NA parameter values are not very intuitive for the actions being performed. It is easier for me to understand and remember the merge syntax: all = TRUE , all.x = TRUE and all.y = TRUE . Since the X[Y] operation resembles merge much more than match , why not use the merge syntax for joins rather than the match function's nomatch parameter? Here are code examples of the 4 join types: # sample X and Y data.tables library(data.table) X <- data.table(t = 1:4, a = (1:4)^2) setkey(X, t) X # t a # 1: 1 1 # 2: 2 4 # 3: 3 9 # 4: 4 16 Y <- data.table(t = 3:6, b = (3:6)^2) setkey(Y, t) Y # t b # 1: 3 9 # 2: 4 16 # 3: 5 25 # 4: 6 36 # all rows from Y - right outer join X[Y] # default # t a b # 1: 3 9 9 # 2: 4 16 16 # 3: 5 NA 25 # 4: 6 NA 36 X[Y, nomatch = NA] # same as above # t a b # 1: 3 9 9 # 2: 4 16 16 #
Tags (comma-separated)
Save Edits
Cancel