missing low surrogate character in surrogate pair, at character offset N (before "…")
This message comes from JSON::PP on Perl 5.038002. In this registry it is produced by 1 distinct mistake. Positions in the real message vary with your input; the causes do not.
Causes
What produces this message
Lone surrogate in a JSON string
Input
{"a": "\uD800"}Exact message: missing low surrogate character in surrogate pair, at character offset 15 (before "(end of string)")
Fix: Never slice strings on UTF-16 boundaries. Slice by code point, or by grapheme cluster if you care about emoji sequences.
Same bug, other languages
What other parsers say about the same input
If a colleague reports a different message for what looks like the same file, this is why.
| Parser | Message for the same inputs |
|---|---|
| Ruby JSON.parse | incomplete surrogate pair at '…' |
| json_decode | Single unpaired UTF-16 surrogate in unicode escape |
| orjson | no low surrogate in string: line L column C (char N) |
| serde_json::from_str | unexpected end of hex escape at line L column C |
Sources
Standards this registry is checked against
- RFC 8259 (STD 90) — The JavaScript Object Notation (JSON) Data Interchange Format, T. Bray, Ed., 2017. The IETF Internet Standard for JSON. Obsoletes RFC 7159 and RFC 4627.
- ECMA-404, 2nd edition — The JSON Data Interchange Syntax, Ecma International, 2017. Ecma’s grammar for JSON. A normative reference of RFC 8259; the two define the same syntax.
- The Unicode Standard — The Unicode Standard, Core Specification, The Unicode Consortium, current. Normatively referenced by RFC 8259 for the definition of a character.
- IEEE 754-2019 — IEEE Standard for Floating-Point Arithmetic, IEEE, 2019. Referenced by RFC 8259 §6 as the basis for its interoperability guidance on numbers.
Only standards bodies and peer-reviewed venues are cited. Error strings on this site are observed by executing each parser; the standards above define what the parser is reacting to.