R DISCUSSION

My R for loop that appends to a vector is very slow: are sapply and vapply actually faster?

Started by neville R loop performancepreallocationapply familyvapply vs sapplyvectorization
5 replies 248 views 6 participants
Latest activity · 30 Sep 2026

My R for loop that appends to a vector is very slow: are sapply and vapply actually faster?

neville R Forum
#1

I process 100,000 rows in a for loop and collect the results with out <- c(out, value); a second loop builds a data frame with rbind one row at a time. Both take minutes. People keep saying "loops are slow in R, use the apply family".

Is the loop itself the problem? And what is the practical difference between lapply, sapply, vapply and apply when I rewrite this?

Community replies 5

Re: My R for loop that appends to a vector is very slow: are sapply and vapply actually faster?

#2

The loop is not the main cost; growing the result is. out <- c(out, value) allocates a new vector one element longer and copies the old one into it on every iteration. Over n iterations that is about n²/2 element copies: for n = 100,000, roughly 5 × 10^9. rbind on a data frame is worse, because each call copies every column.

Preallocate and fill by index: out <- numeric(n), then out[i] <- value. That alone removes the quadratic cost, and the loop then scales linearly with n.

Re: My R for loop that appends to a vector is very slow: are sapply and vapply actually faster?

#3

The apply functions differ in what they return. lapply(x, f) always returns a list with one element per input. sapply tries to simplify that list to a vector or matrix, which is convenient interactively but unpredictable in code: if f returns different lengths, or the input is empty, you get a list instead of a vector.

vapply(x, f, numeric(1)) makes you declare the type and length of each result and throws an error if any call returns something else, so it is the one to use in scripts and functions.

Re: My R for loop that appends to a vector is very slow: are sapply and vapply actually faster?

#4

apply(m, 1, f) is for matrices: it calls f on each row (margin 1) or column (margin 2). Be careful using it on a data frame, because the data frame is first converted to a matrix, and if any column is character, every value becomes a string. For per-row work on mixed-type data, loop over seq_len(nrow(df)) or use Map over the columns.

None of the apply functions is magically faster than a preallocated for loop. They are loops too; they simply handle the allocation of the result for you.

Re: My R for loop that appends to a vector is very slow: are sapply and vapply actually faster?

#5

The real speed comes from vectorised operations, where the loop runs in compiled code. df$p <- df$v * df$i computes all 100,000 rows in one call. rowSums, colMeans, cumsum, diff, ifelse and logical indexing such as x[x > 3.3] cover a large share of what people write element-by-element loops for. For group summaries, tapply(df$v, df$channel, mean) or aggregate replaces a loop over groups.

Before rewriting, measure with system.time() so you know which part is slow.

Re: My R for loop that appends to a vector is very slow: are sapply and vapply actually faster?

#6

For the data frame case, build the pieces first and combine once. Either fill preallocated column vectors and call data.frame() at the end, which is the faster of the two, or collect one small data frame per iteration in a list and finish with do.call(rbind, pieces). Either way the data is combined once instead of being copied afresh on every iteration.

A small robustness point while you are there: write for (i in seq_along(x)) rather than for (i in 1:length(x)). When x is empty, 1:0 is the vector 1, 0 and the loop body runs twice with invalid indices.

TEP COMMUNITY