2026-10-01 11:10:29

A common way to store JSON data is to write one document per line. We call it NDJSON or JSON Lines. Log files, database exports and machine-learning datasets often come in this format. The files can be large, so we may compress them.
{"id":1,"active":true,"user":{"name":"user_1","tags":["guest"]},"score":-3862,"note":"..."}
{"id":2,"active":false,"user":{"name":"user_2","tags":["staff","admin"]},"score":8123,"note":"..."}
How fast can you read such a file? I wrote a small demo with simdjson. My test file has 5 million records: 812 MB of NDJSON. For each record, I read a few fields. I count the active records, and I sum the scores of the active records whose user has the admin tag.
You could decompress the whole file, but it might be better to process the compressed file. So you decompress a chunk, you parse the complete lines in that chunk, and you move the incomplete last line to the front of the buffer. In NDJSON, a newline can never appear inside a document: a newline in a string must be escaped as n. So you can always cut the buffer right after its last newline character. The simdjson library has a function for many documents in one buffer (iterate_many):
simdjson::ondemand::parser parser;
while (true) {
// fill the buffer with decompressed bytes...
size_t cut = eof ? len : last_newline(buf, len) + 1;
simdjson::ondemand::document_stream stream;
parser.iterate_many(buf, cut, cut).get(stream);
for (auto doc : stream) {
accumulate(doc.value_unsafe(), result);
}
if (eof) { break; }
std::memmove(buf, buf + cut, len - cut);
len -= cut;
}
I ran my benchmarks on an Intel Xeon Gold 6548N server (Emerald Rapids) with two sockets, 64 cores and 128 threads. I use GCC 14.
A gzip file is one long compressed stream. The decompressor needs the previous 32 KiB of output to decode what comes next. So you must decompress the file from the start, with one thread. I can still parse in other threads: one thread decompresses chunks and puts them in a queue, and the other threads parse them. It doubles the speed to 2.5 GB/s.
Other formats do better. A zstd or an lz4 file can be made of many independent frames, one after the other. It is still a regular .zst or .lz4 file: the usual command-line tools decompress it as usual.
My program writes a new frame every 256 KiB of JSON, always after a newline. Each frame stores its decompressed size and a checksum. To find where a frame ends, you only need to read a few block headers. You do not need to decompress anything. So the threads can take frames one by one, and each thread decompresses and parses its own frames.

With one thread, zstd and lz4 are no faster than gzip: about 1.1 GB/s. With 64 threads, I get 40 GB/s with zstd and 34 GB/s with lz4. That is 16 times faster than the best I can do with gzip.
The files are not much larger. Small frames compress a bit worse, since each frame starts from scratch, but the zstd file is only 6% larger than the gzip file.
| file | size |
|---|---|
| NDJSON | 812.0 MB |
| gzip (one stream) | 56.9 MB |
| zstd (256 KiB frames) | 60.6 MB |
| lz4 (256 KiB frames) | 108.5 MB |
My data is synthetic. The records are very repetitive, which makes decompression fast. Your data might decompress more slowly.
If you control how your JSON files are written, consider using zstd with many frames. You get files about as small as with gzip, and you can read them many times faster with multiple threads.
Source code: https://github.com/simdjson/simdjson_compressed_demo. The benchmark numbers and the plotting script are in my blog repository.
2026-09-30 20:33:27

Our software represents strings using the UTF-16 or the UTF-8 formats. Most text on the web is UTF-8, but Java, C# or JavaScript represents the strings as UTF-16 to the programmer.
Sometimes we need to transcode (convert) strings. Your browser probably uses the simdutf library for validating or transcoding. It is part of the widely used V8 JavaScript engine.
Up until a few days ago, the simdutf library missed one key feature: if the UTF-8 input is invalid, it did not know how to transcode it to UTF-16. The objective is to replace ill-formed UTF-8 sequences with the replacement character U+FFFD. It is the character that shows up as a weird question mark sometimes. You may have seen it in a broken web site.
The new function convert_utf8_to_utf16_with_replacement always succeeds. It must be used in conjunction with the function utf16_length_from_utf8_with_replacement, which scans and validates the input, recording the offsets of errors if there are some. The usage is as follows.
const char source[] = {'c', 'a', 'f', '\xff'};
size_t length = 4;
simdutf::utf8_to_utf16_result res =
simdutf::utf16_length_from_utf8_with_replacement(source, length);
std::unique_ptr<char16_t[]> utf16{new char16_t[res.count]};
size_t written = simdutf::convert_utf8_to_utf16_with_replacement(
source, length, utf16.get(), res);
Because we have a list of errors before even beginning the transcoding, the bytes between the recorded offsets are valid UTF-8, so we can transcode faster. Each recorded error becomes one U+FFFD.
I timed a scalar decoder against the library on Japanese Wikipedia. The file is 531 KB. I repeated it so the timed buffer is 1.1 MB. I use one core of a Xeon Gold 6548N. GCC 14.3.1 at -O3, best of eight runs.

With no errors, simdutf reaches 7.0 GB/s and the scalar loop 1.7 GB/s.
On valid input, the new approach may even be marginally faster than the previous method, because our initial validation pass allows us to transcode faster afterward.

Compiler: GCC 14.3.1, -O3, one core (taskset -c 2), simdutf icelake kernel. The functions are on the branch utf8-to-utf16-with-replacement.
Credit: The most non-trivial part of this routine was coded by Benjamin Bucher over the summer.
2026-09-30 09:19:54

2026-09-28 20:38:25

The simdjson library is a C++ library to parse and generate JSON. It is used in Node.js, ClickHouse, Meta Velox, StarRocks, Apache Doris, the Ladybird browser and many other systems. We released version 4.0 a year ago, in September 2025. Its headline feature was C++26 static reflection: you could turn a C++ structure into JSON, and back, without writing any glue code.
Today we are releasing version 5.0.
Static reflection is no longer guarded and is an officially supported feature. When your compiler has reflection enabled (e.g., g++ -std=c++26 -freflection with GCC 16), simdjson detects it by itself and the reflection-based functions become available.
When deserializing a C++ structure through reflection, simdjson now uses key selectors (see below) by default: it reads the object in a single pass, whatever the order of the keys. You can return to the previous approach (one lookup per member) with -DSIMDJSON_DISABLE_KEY_SELECTOR_REFLECTION=1.
A positive integer in [2^64, 10^20) is now reported as a big integer (BIGINT_NUMBER), like other overflowing integers, instead of a malformed number. When big integers are parsed as strings, a token such as 123456789123456789123x is now rejected.
One major change is the key selectors. A common task is to extract a few fields from a JSON object. In simdjson 5.0 (C++20 or better), you can name the keys at compile time and visit the object once:
using namespace simdjson;
auto json = R"({ "name": "Daniel", "age": 42, "city": "Montreal" })"_padded;
ondemand::parser parser;
auto doc = parser.iterate(json);
std::string_view name, city;
uint64_t age = 0;
auto result = doc.get_object().for_each<"name", "city", "age">(name, city, age);
// name == "Daniel", city == "Montreal", age == 42
The keys can appear in any order. At compile time, simdjson builds a perfect hash function for your set of keys. At run time, recognizing a key takes a hash computed from a couple of bytes and one comparison. You can also pass one callback per key instead of variables. The iteration stops as soon as all keys have been found.
We added annotations for the data structures for automated (C++26) serialization and deserialization: rename, rename_all, alias, skip, default_value, flatten, deny_unknown_fields, transparent, and so forth.
struct [[= simdjson::rename_all<simdjson::case_style::camel_case>]] User {
std::string first_name;
int64_t user_id;
[[= simdjson::rename<"KEY">]] int api_key;
};
// {"firstName":"Ann","userId":7,"KEY":8}
We often get many JSON documents in one file or one network message. simdjson has long supported streams of documents separated by white space (NDJSON). In 5.0 we added:
stream_format::newline_delimited: you promise that each document sits on its own line. When you only read part of a document, simdjson jumps to the next line instead of walking over the rest of the document.simdjson::slice_at, which cuts a stream into blocks at document boundaries so you can parse the blocks on as many threads as you like. The built-in threaded mode uses at most two threads.We also fixed several bugs in document_stream, found in an audit by Francisco Geiman Thiesen.
There are many other smaller features.
NaN or Infinity, but many systems produce them anyway. If you define SIMDJSON_ENABLE_NAN_INF, simdjson parses them, and serializes them.get_uint8(), get_int8(), get_uint16(), get_int16() check the range for you. With C++23, get_float32() and get_float64() return std::float32_t and std::float64_t. The binary32 value is rounded once, directly from the decimal string, not through a double.parser.parse_unpadded(...)). It is slower than the regular function, but it never reads past the end of your buffer and it does not copy your data.simdjson::padded_input adds padding only when it is needed: when your string ends near a page boundary.std::views::transform and other adaptors.rbegin(), rend()) with no allocation.get_current_position() and revert_position(): if you miss an optional field, you can go back to where you were instead of rescanning the whole object.char8_t (u8) variants of the string accessors in C++20.The simdjson 5.0 release improved performance compared to simdjson 4.0 in some key cases. Let me review some of them.
I built both versions with GCC 16.1 (-O3, CMake Release) and ran them on an Intel Xeon Gold 6548N (Emerald Rapids), pinned to one core.
Let us start with DOM parsing of our standard files (GB/s):
| file | 4.0 | 5.0 | speedup |
|---|---|---|---|
| 4.78 | 4.82 | 1.0 | |
| citm_catalog | 4.82 | 4.79 | 1.0 |
| github_events | 5.33 | 5.30 | 1.0 |
| canada | 1.10 | 1.21 | 1.1 |
| marine_ik | 1.25 | 1.38 | 1.1 |
| mesh | 1.17 | 1.26 | 1.1 |
| numbers | 1.13 | 1.41 | 1.3 |
| twitterescaped | 1.59 | 2.85 | 1.8 |
| update-center | 3.96 | 3.84 | 1.0 |
| apache_builds | 4.95 | 4.78 | 1.0 |
Files full of numbers (canada, marine_ik, mesh, numbers) are 8% to 25% faster. And a file full of escaped Unicode characters (twitterescaped) is almost twice as fast: among other changes, we now decode consecutive uXXXX sequences without going back to the string scanner between them.
We also serialize faster. Printing floating-point numbers used to be a bottleneck. We replaced the ancient Grisu2 by Dragonbox, and removed calls to memcpy and memmove from the hot path.
| file | 4.0 | 5.0 | speedup |
|---|---|---|---|
| 0.94 | 0.96 | 1.0 | |
| citm_catalog | 1.07 | 1.08 | 1.0 |
| gsoc-2018 | 1.10 | 1.25 | 1.1 |
| canada | 0.31 | 0.52 | 1.7 |
| marine_ik | 0.28 | 0.37 | 1.3 |
| mesh | 0.34 | 0.47 | 1.4 |
| numbers | 0.32 | 0.49 | 1.6 |
The simdjson library is a community project. Since version 4.6, contributions came from fior512, 吴杨帆, Alecto Irene Perez, Francisco Geiman Thiesen, Max Bachmann, MoonFlowww, Advit Arora, Jaël Champagne Gareau, Taimoor Kiani, Vasily Pelikh, jmestwa-coder, liyinlong, AlbertoFVisconti, Aylin Dmello, Cuda Chen, Ezra Li, Madhurendra Purbay, Makkar, Pastoray, Paul Dreik, Pavel Kruglov, Piotr Kubaj, Yusuf İhsan Görgel, metsw24-max, neil, pratap singh, Vladimir Saraikin, wankun, xaldarof, Riyane El Qoqui, Justin Li and others. Thank you!
2026-09-27 00:53:53

C# strings are UTF-16. Most characters are one 16-bit code unit. Characters outside the basic multilingual plane, emoji included, take two: a high surrogate (U+D800 to U+DBFF) followed by a low surrogate (U+DC00 to U+DFFF). A surrogate with the wrong neighbor, or with none, is ill-formed.
You should never send an ill-formed string to disk or to the network. It is a bad practice.
In JavaScript, we have fast functions to fix strings or check whether they need fixing:
String.prototype.toWellFormed() replaces every lone surrogate with U+FFFD.isWellFormed() reports whether any replacement is needed.I added both functions to my C# library SimdUnicode, in pull request 54. The algorithm is the same that we contributed to the JavaScript engine V8, so Chrome already fixes strings this way.
string s = UTF16.ToWellFormed(input); // same instance, when the input is already well formed
bool ok = UTF16.IsWellFormed(span);
When the input is well formed, ToWellFormed returns it as is. No allocation.
How are strings fixed? Basically, you replace bad inputs by the replacement character U+FFFD.
Our processors have special instructions called SIMD that allow data parallelism: you can compare multiple values at once. Recent x64 processors from AMD and Intel have better data parallelism than ARM chips, although both have powerful instructions.
The conventional approach in C# to repair a string is a function such as the following.
static void Repair(ReadOnlySpan<char> input, Span<char> output)
{
input.CopyTo(output);
int i = NextError(output, 0);
while (i >= 0)
{
output[i] = 'uFFFD';
i = NextError(output, i + 1);
}
}
// Index of the next lone surrogate at or after 'start', or -1 if none.
static int NextError(ReadOnlySpan<char> s, int start)
{
int i = start;
while (true)
{
int k = s.Slice(i).IndexOfAnyInRange('uD800', 'uDFFF');
if (k < 0) return -1;
i += k;
if (char.IsHighSurrogate(s[i]) && i + 1 < s.Length && char.IsLowSurrogate(s[i + 1]))
i += 2; // valid pair, skip it
else
return i;
}
}
In SimdUnicode, I also use data parallelism.
Let me measure.

On the Xeon, with AVX-512, Latin validates at 69 GB/s against 33 GB/s for IndexOfAnyInRange. The Emoji input is well formed, and it is nothing but surrogate pairs. The runtime search drops to 0.4 GB/s. Our check holds 53 GB/s.

Our results are similar on the M4 Max, although a bit less impressive compared to the Intel results.
Validation can return at the first lone surrogate. The buffer form of ToWellFormed writes every code unit, a copy of the input or U+FFFD. When the input is well formed, it is effectively a memory copy. Thus we can compare the performance against a copy.


Roughly speaking, we are consistently about as fast as a copy.
Versions used: .NET SDK 10.0.400 on Linux, 10.0.103 on macOS. Intel Xeon Gold 6548N (Emerald Rapids). Apple M4 Max.
Clausecker, R., & Lemire, D. (2026). Fixing ill-formed UTF-16 strings with SIMD instructions. Software: Practice and Experience. (arXiv)
2026-09-26 10:48:13
