MoreRSS

site iconDaniel LemireModify

Computer science professor at the University of Quebec (TELUQ), open-source hacker, and long-time blogger.
Please copy the RSS to your reader, or quickly subscribe to:

Inoreader Feedly Follow Feedbin Local Reader

Rss preview of Blog of Daniel Lemire

Parsing compressed JSON at 40 GB/s

2026-10-01 11:10:29

Parsing compressed JSON at 40 GB/s

A common way to store JSON data is to write one document per line. We call it NDJSON or JSON Lines. Log files, database exports and machine-learning datasets often come in this format. The files can be large, so we may compress them.

{"id":1,"active":true,"user":{"name":"user_1","tags":["guest"]},"score":-3862,"note":"..."}
{"id":2,"active":false,"user":{"name":"user_2","tags":["staff","admin"]},"score":8123,"note":"..."}

How fast can you read such a file? I wrote a small demo with simdjson. My test file has 5 million records: 812 MB of NDJSON. For each record, I read a few fields. I count the active records, and I sum the scores of the active records whose user has the admin tag.

You could decompress the whole file, but it might be better to process the compressed file. So you decompress a chunk, you parse the complete lines in that chunk, and you move the incomplete last line to the front of the buffer. In NDJSON, a newline can never appear inside a document: a newline in a string must be escaped as n. So you can always cut the buffer right after its last newline character. The simdjson library has a function for many documents in one buffer (iterate_many):

simdjson::ondemand::parser parser;
while (true) {
  // fill the buffer with decompressed bytes...
  size_t cut = eof ? len : last_newline(buf, len) + 1;
  simdjson::ondemand::document_stream stream;
  parser.iterate_many(buf, cut, cut).get(stream);
  for (auto doc : stream) {
    accumulate(doc.value_unsafe(), result);
  }
  if (eof) { break; }
  std::memmove(buf, buf + cut, len - cut);
  len -= cut;
}

I ran my benchmarks on an Intel Xeon Gold 6548N server (Emerald Rapids) with two sockets, 64 cores and 128 threads. I use GCC 14.

A gzip file is one long compressed stream. The decompressor needs the previous 32 KiB of output to decode what comes next. So you must decompress the file from the start, with one thread. I can still parse in other threads: one thread decompresses chunks and puts them in a queue, and the other threads parse them. It doubles the speed to 2.5 GB/s.

Other formats do better. A zstd or an lz4 file can be made of many independent frames, one after the other. It is still a regular .zst or .lz4 file: the usual command-line tools decompress it as usual.

My program writes a new frame every 256 KiB of JSON, always after a newline. Each frame stores its decompressed size and a checksum. To find where a frame ends, you only need to read a few block headers. You do not need to decompress anything. So the threads can take frames one by one, and each thread decompresses and parses its own frames.

Decompressing and parsing NDJSON

With one thread, zstd and lz4 are no faster than gzip: about 1.1 GB/s. With 64 threads, I get 40 GB/s with zstd and 34 GB/s with lz4. That is 16 times faster than the best I can do with gzip.

The files are not much larger. Small frames compress a bit worse, since each frame starts from scratch, but the zstd file is only 6% larger than the gzip file.

file size
NDJSON 812.0 MB
gzip (one stream) 56.9 MB
zstd (256 KiB frames) 60.6 MB
lz4 (256 KiB frames) 108.5 MB

My data is synthetic. The records are very repetitive, which makes decompression fast. Your data might decompress more slowly.

If you control how your JSON files are written, consider using zstd with many frames. You get files about as small as with gzip, and you can read them many times faster with multiple threads.

Source code: https://github.com/simdjson/simdjson_compressed_demo. The benchmark numbers and the plotting script are in my blog repository.

Transcoding UTF-8 to UTF-16 with replacement at gigabytes per second

2026-09-30 20:33:27

Transcoding UTF-8 to UTF-16 with replacement

Our software represents strings using the UTF-16 or the UTF-8 formats. Most text on the web is UTF-8, but Java, C# or JavaScript represents the strings as UTF-16 to the programmer.

Sometimes we need to transcode (convert) strings. Your browser probably uses the simdutf library for validating or transcoding. It is part of the widely used V8 JavaScript engine.

Up until a few days ago, the simdutf library missed one key feature: if the UTF-8 input is invalid, it did not know how to transcode it to UTF-16. The objective is to replace ill-formed UTF-8 sequences with the replacement character U+FFFD. It is the character that shows up as a weird question mark sometimes. You may have seen it in a broken web site.

The new function convert_utf8_to_utf16_with_replacement always succeeds. It must be used in conjunction with the function utf16_length_from_utf8_with_replacement, which scans and validates the input, recording the offsets of errors if there are some. The usage is as follows.

const char source[] = {'c', 'a', 'f', '\xff'};
size_t length = 4;
simdutf::utf8_to_utf16_result res =
    simdutf::utf16_length_from_utf8_with_replacement(source, length);
std::unique_ptr<char16_t[]> utf16{new char16_t[res.count]};
size_t written = simdutf::convert_utf8_to_utf16_with_replacement(
    source, length, utf16.get(), res);

Because we have a list of errors before even beginning the transcoding, the bytes between the recorded offsets are valid UTF-8, so we can transcode faster. Each recorded error becomes one U+FFFD.

I timed a scalar decoder against the library on Japanese Wikipedia. The file is 531 KB. I repeated it so the timed buffer is 1.1 MB. I use one core of a Xeon Gold 6548N. GCC 14.3.1 at -O3, best of eight runs.

UTF-8 to UTF-16 with replacement against a scalar decoder, Japanese Wikipedia

With no errors, simdutf reaches 7.0 GB/s and the scalar loop 1.7 GB/s.

On valid input, the new approach may even be marginally faster than the previous method, because our initial validation pass allows us to transcode faster afterward.

Valid UTF-8 to UTF-16, ordinary conversion and conversion with replacement

Compiler: GCC 14.3.1, -O3, one core (taskset -c 2), simdutf icelake kernel. The functions are on the branch utf8-to-utf16-with-replacement.

Credit: The most non-trivial part of this routine was coded by Benjamin Bucher over the summer.

How do you deal with people who strongly disagree with your views?

2026-09-30 09:19:54

 
1. Begin by genuinely listening to them. Make sure that you understand them. It is often useful to ask for confirmation: restate what you think they are saying and make sure that you understand. Physically, it helps to turn your shoulders so that you face them. Don’t look at your screen while they talk.
 
2. Ask more questions. You can often effortlessly weaken the stance of someone who has not given much thought to an issue just by asking them questions. Don’t ‘challenge them’. Ask genuine questions. The objective is not to break their views with questions, but to establish the playing ground. Many people are unclear about what they think they believe. They may simply be repeated what they heard.
 
3. Focus on what you share in common. Often, the disagreement is smaller than you think. Furthermore, you can sometimes entirely avoid a conflict by establishing strong agreement on many points.
 
4. Unless you are clearly a knowledgeable person, do not immediately express your opinion. Begin by establishing your credibility. Refer to hard facts, data. Refer to people or documents you have consulted. This, in turn, should have been done ahead of time. Never engage in a debate without having done at least a bit of research and reflection.
 
5. Use language that fits the issue. You need to learn the vocabulary. Learn the technical terms and use them, although you may want to remain accessible, so explain them as well.
 
6. Tell stories. People understand by emotions as much as by their reason. Ideally, the story should be true and it should be backed by hard evidence.
 
7. Avoid personal attacks and do not tolerate them against you and others. If someone is insulting you, if it is fair to say so. You may say that you feel offended. It is the one case where you can and maybe should go on the offensive. Don’t answer by an insult, but attack the fact that they insulted you. You may ask for an apology.

simdjson 5.0 is out

2026-09-28 20:38:25

simdjson 5.0 is out

The simdjson library is a C++ library to parse and generate JSON. It is used in Node.js, ClickHouse, Meta Velox, StarRocks, Apache Doris, the Ladybird browser and many other systems. We released version 4.0 a year ago, in September 2025. Its headline feature was C++26 static reflection: you could turn a C++ structure into JSON, and back, without writing any glue code.

Today we are releasing version 5.0.

  1. Static reflection is no longer guarded and is an officially supported feature. When your compiler has reflection enabled (e.g., g++ -std=c++26 -freflection with GCC 16), simdjson detects it by itself and the reflection-based functions become available.

  2. When deserializing a C++ structure through reflection, simdjson now uses key selectors (see below) by default: it reads the object in a single pass, whatever the order of the keys. You can return to the previous approach (one lookup per member) with -DSIMDJSON_DISABLE_KEY_SELECTOR_REFLECTION=1.

  3. A positive integer in [2^64, 10^20) is now reported as a big integer (BIGINT_NUMBER), like other overflowing integers, instead of a malformed number. When big integers are parsed as strings, a token such as 123456789123456789123x is now rejected.

One major change is the key selectors. A common task is to extract a few fields from a JSON object. In simdjson 5.0 (C++20 or better), you can name the keys at compile time and visit the object once:

using namespace simdjson;
auto json = R"({ "name": "Daniel", "age": 42, "city": "Montreal" })"_padded;
ondemand::parser parser;
auto doc = parser.iterate(json);
std::string_view name, city;
uint64_t age = 0;
auto result = doc.get_object().for_each<"name", "city", "age">(name, city, age);
// name == "Daniel", city == "Montreal", age == 42

The keys can appear in any order. At compile time, simdjson builds a perfect hash function for your set of keys. At run time, recognizing a key takes a hash computed from a couple of bytes and one comparison. You can also pass one callback per key instead of variables. The iteration stops as soon as all keys have been found.

We added annotations for the data structures for automated (C++26) serialization and deserialization: rename, rename_all, alias, skip, default_value, flatten, deny_unknown_fields, transparent, and so forth.

struct [[= simdjson::rename_all<simdjson::case_style::camel_case>]] User {
  std::string first_name;
  int64_t user_id;
  [[= simdjson::rename<"KEY">]] int api_key;
};
// {"firstName":"Ann","userId":7,"KEY":8}

We often get many JSON documents in one file or one network message. simdjson has long supported streams of documents separated by white space (NDJSON). In 5.0 we added:

  • RFC 7464 JSON text sequences (each document preceded by the record separator character) and comma-separated documents;
  • stream_format::newline_delimited: you promise that each document sits on its own line. When you only read part of a document, simdjson jumps to the next line instead of walking over the rest of the document.
  • simdjson::slice_at, which cuts a stream into blocks at document boundaries so you can parse the blocks on as many threads as you like. The built-in threaded mode uses at most two threads.

We also fixed several bugs in document_stream, found in an audit by Francisco Geiman Thiesen.

There are many other smaller features.

  • NaN and infinity: JSON does not allow NaN or Infinity, but many systems produce them anyway. If you define SIMDJSON_ENABLE_NAN_INF, simdjson parses them, and serializes them.
  • Narrow types: get_uint8(), get_int8(), get_uint16(), get_int16() check the range for you. With C++23, get_float32() and get_float64() return std::float32_t and std::float64_t. The binary32 value is rounded once, directly from the decimal string, not through a double.
  • The DOM API can parse a buffer that has no padding (parser.parse_unpadded(...)). It is slower than the regular function, but it never reads past the end of your buffer and it does not copy your data.
  • With C++17, simdjson::padded_input adds padding only when it is needed: when your string ends near a page boundary.
  • C++20 ranges: you can pipe On-Demand arrays and objects into std::views::transform and other adaptors.
  • DOM arrays support reverse iteration (rbegin(), rend()) with no allocation.
  • On-Demand objects offer get_current_position() and revert_position(): if you miss an optional field, you can go back to where you were instead of rescanning the whole object.
  • char8_t (u8) variants of the string accessors in C++20.
  • Better pretty printing with the FracturedJson style, including tables.
  • Memory-mapped files under Windows.
  • We support the memory-safe compiler Fil-C.

The simdjson 5.0 release improved performance compared to simdjson 4.0 in some key cases. Let me review some of them.

I built both versions with GCC 16.1 (-O3, CMake Release) and ran them on an Intel Xeon Gold 6548N (Emerald Rapids), pinned to one core.

Let us start with DOM parsing of our standard files (GB/s):

file 4.0 5.0 speedup
twitter 4.78 4.82 1.0
citm_catalog 4.82 4.79 1.0
github_events 5.33 5.30 1.0
canada 1.10 1.21 1.1
marine_ik 1.25 1.38 1.1
mesh 1.17 1.26 1.1
numbers 1.13 1.41 1.3
twitterescaped 1.59 2.85 1.8
update-center 3.96 3.84 1.0
apache_builds 4.95 4.78 1.0

Files full of numbers (canada, marine_ik, mesh, numbers) are 8% to 25% faster. And a file full of escaped Unicode characters (twitterescaped) is almost twice as fast: among other changes, we now decode consecutive uXXXX sequences without going back to the string scanner between them.

We also serialize faster. Printing floating-point numbers used to be a bottleneck. We replaced the ancient Grisu2 by Dragonbox, and removed calls to memcpy and memmove from the hot path.

file 4.0 5.0 speedup
twitter 0.94 0.96 1.0
citm_catalog 1.07 1.08 1.0
gsoc-2018 1.10 1.25 1.1
canada 0.31 0.52 1.7
marine_ik 0.28 0.37 1.3
mesh 0.34 0.47 1.4
numbers 0.32 0.49 1.6

The simdjson library is a community project. Since version 4.6, contributions came from fior512, 吴杨帆, Alecto Irene Perez, Francisco Geiman Thiesen, Max Bachmann, MoonFlowww, Advit Arora, Jaël Champagne Gareau, Taimoor Kiani, Vasily Pelikh, jmestwa-coder, liyinlong, AlbertoFVisconti, Aylin Dmello, Cuda Chen, Ezra Li, Madhurendra Purbay, Makkar, Pastoray, Paul Dreik, Pavel Kruglov, Piotr Kubaj, Yusuf İhsan Görgel, metsw24-max, neil, pratap singh, Vladimir Saraikin, wankun, xaldarof, Riyane El Qoqui, Justin Li and others. Thank you!

How fast can you fix a UTF-16 string in C#?

2026-09-27 00:53:53

How fast can you fix a UTF-16 string in C#

C# strings are UTF-16. Most characters are one 16-bit code unit. Characters outside the basic multilingual plane, emoji included, take two: a high surrogate (U+D800 to U+DBFF) followed by a low surrogate (U+DC00 to U+DFFF). A surrogate with the wrong neighbor, or with none, is ill-formed.

You should never send an ill-formed string to disk or to the network. It is a bad practice.

In JavaScript, we have fast functions to fix strings or check whether they need fixing:

  • String.prototype.toWellFormed() replaces every lone surrogate with U+FFFD.
  • isWellFormed() reports whether any replacement is needed.

I added both functions to my C# library SimdUnicode, in pull request 54. The algorithm is the same that we contributed to the JavaScript engine V8, so Chrome already fixes strings this way.

string s = UTF16.ToWellFormed(input); // same instance, when the input is already well formed
bool ok = UTF16.IsWellFormed(span);

When the input is well formed, ToWellFormed returns it as is. No allocation.

How are strings fixed? Basically, you replace bad inputs by the replacement character U+FFFD.

Our processors have special instructions called SIMD that allow data parallelism: you can compare multiple values at once. Recent x64 processors from AMD and Intel have better data parallelism than ARM chips, although both have powerful instructions.

The conventional approach in C# to repair a string is a function such as the following.

static void Repair(ReadOnlySpan<char> input, Span<char> output)
{
    input.CopyTo(output);
    int i = NextError(output, 0);
    while (i >= 0)
    {
        output[i] = 'uFFFD';
        i = NextError(output, i + 1);
    }
}
// Index of the next lone surrogate at or after 'start', or -1 if none.
static int NextError(ReadOnlySpan<char> s, int start)
{
    int i = start;
    while (true)
    {
        int k = s.Slice(i).IndexOfAnyInRange('uD800', 'uDFFF');
        if (k < 0) return -1;
        i += k;
        if (char.IsHighSurrogate(s[i]) && i + 1 < s.Length && char.IsLowSurrogate(s[i + 1]))
            i += 2; // valid pair, skip it
        else
            return i;
    }
}

In SimdUnicode, I also use data parallelism.

Let me measure.

Is the UTF-16 string well formed? Intel Xeon Gold 6548N

On the Xeon, with AVX-512, Latin validates at 69 GB/s against 33 GB/s for IndexOfAnyInRange. The Emoji input is well formed, and it is nothing but surrogate pairs. The runtime search drops to 0.4 GB/s. Our check holds 53 GB/s.

Is the UTF-16 string well formed? Apple M4 Max

Our results are similar on the M4 Max, although a bit less impressive compared to the Intel results.

Validation can return at the first lone surrogate. The buffer form of ToWellFormed writes every code unit, a copy of the input or U+FFFD. When the input is well formed, it is effectively a memory copy. Thus we can compare the performance against a copy.

Copy the string, replace lone surrogates. Intel Xeon Gold 6548N

Copy the string, replace lone surrogates. Apple M4 Max

Roughly speaking, we are consistently about as fast as a copy.

Versions used: .NET SDK 10.0.400 on Linux, 10.0.103 on macOS. Intel Xeon Gold 6548N (Emerald Rapids). Apple M4 Max.

Clausecker, R., & Lemire, D. (2026). Fixing ill-formed UTF-16 strings with SIMD instructions. Software: Practice and Experience. (arXiv)

Source code.

A thesis isn’t enough for a PhD

2026-09-26 10:48:13

To get a PhD, you typically have to enroll in a graduate program. Then you complete a few relatively easy courses.
 
You might need to pass a comprehensive examination that checks whether you have a basic understanding of the field. A few students fail at this point, but not many.Then you write a thesis and defend it.
 
In theory, you could write a strong thesis and still fail the oral defense. That is uncommon. The defense is typically a public event, and there may be guests. Failing a student at that stage would be a public humiliation. I have seen students who could not answer basic questions still receive their PhDs because the thesis itself was good enough.
 
Thus, to a very good approximation, getting a PhD amounted to writing a thesis.But we have a problem. AI can write something that looks quite a bit like a thesis. With a bit of prompting, you can produce something that looks like a PhD thesis.I have been telling everyone I can that we have a big problem.Apparently other people have realized this as well. People at Harvard are now saying that a thesis is insufficient for a PhD.
 
 
 
« The dissertation has in the past frequently been used as a proxy for the kind of mathematical development of a student that we expect: acquiring mathematical knowledge and demonstrating independent achievement. However, with modern AI systems, dissertations are not (…) a reliable tool for evaluating students, and PhDs should not be awarded primarily on the basis of the text of the dissertation. We thus recommend regular, multi-faceted evaluations of students in person, by multiple faculty, as the central component of assessment. These evaluations must be rigorous, not pro-forma, and should produce reports that might become part of the student’s dossier for future employment. »