MoreRSS

site iconDaniel LemireModify

Computer science professor at the University of Quebec (TELUQ), open-source hacker, and long-time blogger.
Please copy the RSS to your reader, or quickly subscribe to:

Inoreader Feedly Follow Feedbin Local Reader

Rss preview of Blog of Daniel Lemire

Faster Python startup with lazy imports

2026-10-09 11:15:06

When a Python program starts, it needs to load all its dependencies (import). With the upcoming version of Python, it is possible to use lazy imports instead.

lazy import json
lazy from decimal import Decimal

In this instance, the name json is no longer the module, but a mere placeholder object. The actual import only happens when you use json. If you never use json, the module is never loaded.

How much does it help? Let us measure.

I wrote a toy command-line tool. It imports sixteen modules: json, csv, decimal, sqlite3, asyncio, email.parser, http.client, urllib.request, xml.etree.ElementTree, zipfile, tarfile, statistics and some popular third-party packages (numpy, pandas, requests, rich). The tool has three paths:

  • --version prints a string and needs nothing;
  • mean reads a small CSV file with the csv module and calls statistics.fmean;
  • stats loads the same file with pandas and prints a summary.

The lazy version of the tool is identical, except that I prefix the imports with lazy. It is a one-word change per line.

I use Python 3.15 (3.15.0b4) on an Intel Xeon Gold 6548N (Emerald Rapids), with numpy 2.5, pandas 3.0 and requests 2.34.

command eager imports lazy imports speedup
--version 295 ms 20 ms 15×
mean 296 ms 24 ms 12×
stats 299 ms 224 ms 1.3×

Eager imports take about 300 ms. Lazy imports take 20 ms for --version, 24 ms for mean, and 224 ms for stats.

Printing the version number goes from 295 ms to 20 ms. If you subtract the interpreter startup (11.7 ms), the cost of the imports goes from 283 ms to about 8 ms: a 35-fold reduction. The mean command, which needs a few standard modules, is twelve times faster.

There is a downside: errors move. With an eager import, a missing module fails at startup. With a lazy import, it fails at the first use, perhaps deep inside a function, perhaps hours later in a long-running server.

My source code is available.

Mastering SIMD with Java Vector API

2026-10-07 10:07:05

Cover of Mastering SIMD with Java Vector API by Roman Snytsar

I write a lot about data parallelism. About how our processors have special instructions that can do several operations at once. They are often critical to get the best performance out of a CPU.

Compilers are sometimes able to take advantage of these instructions. But it is hard to tell when it will work.

For Java programmers, you can write directly for data parallelism since JDK 16. With enough skill, you might multiply the performance of a few key functions.

The Vector API is in the incubator module jdk.incubator.vector. It still not officially supported, but you can enable it by passing --add-modules=jdk.incubator.vector. I expect that it will soon become a mainstream feature in Java.

Roman Snytsar wrote a book about it: Mastering SIMD with Java Vector API, and the subtitle is Unlocking Single-Core Performance Through SIMD Optimization. I had the honor of being the technical reviewer. I learned a few things while reading the book and I enjoyed it.

Who has time for books, especially technical books?

What a book offers is time to reflect. Reading technical books today is probably just as relevant as it ever was. The author takes you on a story… and gets you to think. 

Cover of Mastering SIMD with Java Vector API by Roman Snytsar

Snytsar’s book works from  problems like sum an array, compute a mean and a standard deviation, remove duplicates, merge two sorted arrays, and so forth. A lot of them are similar to the problems you will encounter as a programmer.  He starts from the ordinary solution and rebuilds it with the Java Vector API. He then runs benchmarks. He looks at the assembly code.

The book covers performance issues such as unrolling, asymmetric loops, structural hazards, dependencies. Snytsar shows us that, in some instances, data parallelism can fail to bear fruits. The negative lessons are just as important as the positive ones.

The Java Vector API is a thick layer of abstractions, but Snytsar shows that some tricks work better on some hardware than others.  I liked the chapter on removing duplicates, which is built on compress. It is one instance where the specific hardware matters. And you are unlikely to just find out about it yourself.

It is a book worth buying if you are a Java programmer with a focus on performance. 

The book: Roman Snytsar, Mastering SIMD with Java Vector API, Apress, 2026.

Faster software linking with mold

2026-10-06 16:00:20

When you build a program, the compiler turns each source file into an object file. Then a linker stitches all the object files and libraries into one executable.

On Linux, the default linker is usually GNU ld (also called bfd). The GNU binutils also ship gold, an alternative linker that was designed to be faster. Gold was built a Google and first made available in 2008.

The mold linker is a newer linker written by Rui Ueyama, it was first released in 2021. It is meant as a drop-in replacement for GNU ld, and its main selling point is speed. The latest release is version 3.

How much faster is it on a large project? I built Node.js from its main branch on an Intel Xeon Gold 6548N server (Emerald Rapids, two sockets, 64 cores and 128 threads) running Linux with GCC 14.3. The final node executable weighs about 160 MB.

I captured the command that links the node executable and ran it six times with each linker. I report the median.

linker time to link node
GNU ld (bfd) 2.41 2.52 s
GNU gold 2.41 1.46 s
mold 3.0.0, 1 thread 0.49 s
mold 3.0.0, 8 threads 0.13 s
mold 3.0.0, 128 threads 0.11 s

With multithreading, mold is about 24 times faster than GNU ld and about 14 times faster than gold.

Even if I restrict mold to a single thread, it links Node.js in half a second, five times faster than GNU ld. With only 8 threads, you get nearly all of the benefit.

Using mold is easy. You can pass -fuse-ld=mold to GCC or clang. Or you can wrap your whole build:

mold -run make -j128

The -run option intercepts every call to the default linker and redirects it to mold. You do not need to change the build scripts.

Of course, two seconds saved on a link does not matter much if you build Node.js once. But a developer who edits a file and rebuilds dozens of times a day pays the link cost each time. With mold, the link step becomes nearly free.

My scripts and raw results are available.

Ephemeral testing

2026-10-05 16:00:51

We have many ways to ensure software quality. Unit testing. Fuzz testing. Integration testing. And so forth.
 
I’d like to propose a method that was unthinkable before: ephemeral testing. (Ephemeral is a fancy word for ‘throw away’ or ‘temporary’.)
 
You write your code. You build your software component. Or the AI agent does it for you, it does not matter.
 
Then you ask an AI agent to build on it: an application, another layer, maybe several. You have it test what it built. You do not assess the original work directly. You assess how good the software built on top of it is.
 
It is a form of integration testing. The difference is that the software on top is entirely ephemeral. You throw it away when you are done.
 
A library with a clean API, stable invariants, and useful errors lets the agent produce something that works quickly. A library with hidden state, surprising defaults, or incomplete docs produces a pile of patches and failures. The failures are evidence about your code, not about the agent.
 
You can repeat it. Different agents, different tasks, same foundation.
 
In effect, instead of building the core while trying to anticipate what might be needed at the other layers, you just simulate the other layers by actually building them.
 
Of course, you could argue that with AI, you can rebuild everything whenever you need to. But that’s not practical. You need some form of stability.
 
I have been applying this trick to various projects. As I consider a new feature, I ask my AI to prototype quickly what I might later build based on what I am doing it. Ephemeral testing works for me thus far.

Research paper overload: submissions capped at two a month

2026-10-05 01:07:27

Much of my research is on arXiv. I was one of the early adopters. It is simply a repository of research papers with PDF and metadata.

It makes it convenient to find research without any paywall.

Physicists started it in 1991. Math and computer science followed gradually.

It relies on volunteers to moderate it, because we don’t want garbage or spam.

On October 1, they capped the submissions to two per person per month.

Why?

arXiv got over 40,000 submissions in September.

Two years ago, it was about 20,000.

Doubling every two years. It is unsustainable for human beings.

Papers posted to arXiv each month

In related news, Google stopped taking new bug reports in its open-source bounty program. They could not cope with the influx.

The writing is on the wall.

Google’s first orbital data center is in orbit

2026-10-05 00:49:29

Google’s first orbital data center is in orbit (October 1st). They say it is about the size of a fridge.

It has one kilowatt of solar power. It is about the power of a household.

It was launched on a Falcon 9 rocket (SpaceX).

It should run for about a year. Google plans to launch two more orbital data centres next year. The plan is to use laser communications.

For people complaining that the latency is going to be a problem with orbital data centres…

An orbital data centre can be effectively in your line of sight and it can be relatively close (much closer than the diameter of the USA).

Orbital latency compared with California to New York


Further if you have a cluster of space data centres, they can remain close and in a line of sight.

So you are not going to run your video games or do high frequency trading from space, but for compute, latency is unlikely to be the bottleneck.

The great features are space (no need to wait five years for a permit) and power (no need to negotiate with a local government for electricity). Solar power in space is great because it runs 24h a day.