Python DISCUSSION

Python 3 serial data arrives as b'...' bytes: how do I convert bytes to str correctly?

Started by reska bytes vs strdecode and encodepyserialUnicodeDecodeErrorstruct unpack
5 replies 248 views 6 participants
Latest activity · 30 Sep 2026

Python 3 serial data arrives as b'...' bytes: how do I convert bytes to str correctly?

reska Python Forum
#1

Reading a line from an Arduino with pyserial gives me b'23.5\r\n' instead of 23.5. Comparing it with a string fails, line == "23.5" is always False, and "Temp: " + line raises TypeError: can only concatenate str (not "bytes") to str. Wrapping it in str(line) gives the text b'23.5\r\n' with the b and the quotes included.

What is the difference between bytes and str in Python 3, how do I convert properly, and what should I do when the device sometimes sends garbage bytes?

Community replies 5

Re: Python 3 serial data arrives as b'...' bytes: how do I convert bytes to str correctly?

#2

str is a sequence of Unicode characters. bytes is a sequence of integers from 0 to 255, which is what actually travels over a wire or sits in a file. An encoding maps between the two. Serial ports, sockets and files opened in binary mode deliver bytes, because Python cannot know what the device meant by them.

To get text, decode: text = line.decode("ascii") or "utf-8". To send text, encode: ser.write("START\n".encode("ascii")), or use a bytes literal, ser.write(b"START\n"). str(line) does not decode; without an encoding argument it returns the printable representation, which is why you see the b and the quotes.

Re: Python 3 serial data arrives as b'...' bytes: how do I convert bytes to str correctly?

#3

For your reading: value = float(line.decode("ascii").strip()). The strip() removes the trailing \r\n and any other surrounding whitespace.

Indexing differs between the two types. line[0] on bytes gives an integer, 50 for the character 2, while the slice line[0:1] gives the bytes object b'2'. Comparing bytes with str never raises an error; it is simply unequal every time, which is why such tests fail silently. Either decode first, or compare with a bytes literal: if line.startswith(b"OK"):.

Re: Python 3 serial data arrives as b'...' bytes: how do I convert bytes to str correctly?

#4

On garbage: decode raises UnicodeDecodeError for a byte that is invalid in the chosen encoding, for example 0xFF in ASCII. With a noisy line, or a device that resets in the middle of a message, that will happen sooner or later, so pick a policy. line.decode("ascii", errors="replace") substitutes a replacement character for bad bytes and errors="ignore" drops them. For measurements, catching the exception and discarding the whole line is usually right, because a line with a corrupted byte should not be trusted.

Frequent garbage points to a cause worth fixing: a baud rate mismatch, or reading before the board has finished resetting. Many Arduino boards reset when the port is opened, so discard the first partial line.

Re: Python 3 serial data arrives as b'...' bytes: how do I convert bytes to str correctly?

#5

If the device sends binary data rather than text, do not decode at all. Use the struct module: struct.unpack("<Hf", payload) reads a little-endian unsigned 16-bit integer followed by a 32-bit float from 6 bytes. int.from_bytes(data[0:2], "little") converts a single integer, and data.hex() prints bytes for debugging, for example b'\x01\xff'.hex() gives '01ff'.

Always state the byte order with < or >. In the default native mode struct also inserts alignment padding, so the format "Hf" expects 8 bytes on common platforms instead of 6, and the unpack fails or reads the wrong bytes.

Re: Python 3 serial data arrives as b'...' bytes: how do I convert bytes to str correctly?

#6

For text files, let the I/O layer decode: open(path, encoding="utf-8") returns str lines. Always pass encoding, because the default depends on the platform and may be a legacy code page on Windows.

Be careful with multi-byte characters on a stream. The degree sign is two bytes in UTF-8 (0xC2 0xB0), and it can be split across two read() calls, in which case decoding each chunk separately fails. Decode only complete lines, or use an incremental decoder from the codecs module. If a device sends the degree sign as the single byte 0xB0, it is using Latin-1 rather than UTF-8; decode with "latin-1", which maps every byte value to a character and therefore never raises.

TEP COMMUNITY