R DISCUSSION

R subsetting: when do I use [ ], [[ ]] or $ on a list or data frame?

Started by ayesha R subsettingsingle vs double bracketsdollar operatordata frame columnslist extraction
5 replies 248 views 6 participants
Latest activity · 30 Sep 2026

R subsetting: when do I use [ ], [[ ]] or $ on a list or data frame?

ayesha R Forum
#1

I have a data frame of sensor logs and a list of fitted models. mean(df["temp"]) returns NA with a warning, while mean(df[["temp"]]) and mean(df$temp) work. With my list, summary(models[1]) prints a short table about a list instead of the model summary I get from summary(models[[1]]).

What is the rule behind single brackets, double brackets and the dollar sign, and how do I pick the right one when the column name is stored in a variable?

Community replies 5

Re: R subsetting: when do I use [ ], [[ ]] or $ on a list or data frame?

#2

Single brackets return a subset of the same kind of container; double brackets extract one element out of it. For a list, models[1] is a list of length one that still contains the model, and models[[1]] is the model itself. A data frame is a list of columns, so df["temp"] is a one-column data frame and df[["temp"]] is the numeric vector. mean() needs a vector, which is why the first form gives NA with the "argument is not numeric or logical" warning.

Re: R subsetting: when do I use [ ], [[ ]] or $ on a list or data frame?

#3

$ is shorthand for [[ with a literal name: df$temp is df[["temp"]]. The name after the dollar sign is not evaluated, so with col <- "temp", df$col looks for a column literally called col and returns NULL. Use df[[col]] when the name is in a variable.

On lists and base data frames $ also does partial matching: df$te finds temp if it is the only name starting that way. It is convenient at the console and a source of silent bugs in scripts, so write full names in code.

Re: R subsetting: when do I use [ ], [[ ]] or $ on a list or data frame?

#4

The two-index form has its own trap. df[, "temp"] on a base data frame drops to a vector because drop = TRUE is the default when a single column is selected, but df[, c("temp", "rh")] stays a data frame. Code that selects a variable number of columns then breaks when the selection happens to have length one.

Write df[, cols, drop = FALSE] when you always want a data frame back. Tibbles behave differently here: [ on a tibble always returns a tibble.

Re: R subsetting: when do I use [ ], [[ ]] or $ on a list or data frame?

#5

They also differ on bad indices, which is useful for catching mistakes. On a vector, x[10] beyond the end returns NA quietly, while x[[10]] stops with a "subscript out of bounds" error. On a list, a name that does not exist gives NULL with both lst[["missing"]] and lst$missing.

[[ is meant for exactly one element, so it cannot silently hand back several elements or none. In functions and loops, [[ is the safer default whenever you mean one element.

Re: R subsetting: when do I use [ ], [[ ]] or $ on a list or data frame?

#6

Assignment follows the same logic. lst[["a"]] <- NULL removes the element a; to store an actual NULL you need lst["a"] <- list(NULL). And df$ratio <- df$v / df$i adds a column.

A quick way to see what you have is str(): str(df["temp"]) reports a data.frame with one variable, str(df[["temp"]]) reports a plain num vector. Checking with str or class before passing something to a function settles most of these questions in a few seconds.

TEP COMMUNITY