R DISCUSSION

Why does x == NA never work in R, and how should missing values be tested and removed?

Started by amine R missing valuesis.nana.rmNA comparisoncomplete.cases
4 replies 248 views 5 participants
Latest activity · 30 Sep 2026

Why does x == NA never work in R, and how should missing values be tested and removed?

amine R Forum
#1

My logger writes NA when a sensor read fails. In R, df[df$temp == NA, ] returns rows full of NA instead of the rows with missing temperature, mean(df$temp) is NA, and if (x > 30) stops with "missing value where TRUE/FALSE needed" on some rows.

What is the logic behind these results, and what are the correct ways to detect, skip or replace missing values?

Community replies 4

Re: Why does x == NA never work in R, and how should missing values be tested and removed?

#2

NA means "value unknown", and any comparison with an unknown value is itself unknown. 5 == NA is NA, and so is NA == NA, because two unknown values might or might not be equal. So df$temp == NA is a vector of nothing but NA, and indexing rows with an NA index gives a row of NAs.

The test for missingness is is.na(x): df[is.na(df$temp), ] gives the rows with a missing temperature and df[!is.na(df$temp), ] the rest.

Re: Why does x == NA never work in R, and how should missing values be tested and removed?

#3

The same propagation explains the other two symptoms. Arithmetic with NA gives NA, so one missing sample makes mean, sum, max and sd return NA. Most summary functions take na.rm = TRUE to drop missing values first: mean(df$temp, na.rm = TRUE).

if needs a single TRUE or FALSE, and NA > 30 is neither, hence the error. Guard it: if (!is.na(x) && x > 30). The && short-circuits, so the comparison is never evaluated for a missing value.

Re: Why does x == NA never work in R, and how should missing values be tested and removed?

#4

Logical subsetting has a quieter version of the first problem. df[df$temp > 30, ] keeps the rows where the condition is TRUE and also returns an all-NA row for every row where temp is missing. Wrap the condition in which(), which returns only the positions that are TRUE, or use subset(df, temp > 30), which drops NA conditions.

To count the gaps, sum(is.na(df$temp)) works for one column and colSums(is.na(df)) gives the count for every column.

Re: Why does x == NA never work in R, and how should missing values be tested and removed?

#5

For removal and replacement: na.omit(df) or df[complete.cases(df), ] drops every row that has a missing value in any column, which can discard a lot of good data if only one channel failed; pass a subset of columns to complete.cases to be selective.

x[is.na(x)] <- 0 replaces them, but think about whether zero is honest for a measurement. A failed temperature read replaced by 0 will pull the mean down and looks like real data in a plot. Leaving the gap as NA and using na.rm = TRUE is usually the more truthful choice. Also make sure the file reader recognises your marker; other spellings need na.strings.

TEP COMMUNITY