2026-10-09 11:15:06

When a Python program starts, it needs to load all its dependencies (import). With the upcoming version of Python, it is possible to use lazy imports instead.
lazy import json
lazy from decimal import Decimal
In this instance, the name json is no longer the module, but a mere placeholder object. The actual import only happens when you use json. If you never use json, the module is never loaded.
How much does it help? Let us measure.
I wrote a toy command-line tool. It imports sixteen modules: json, csv, decimal, sqlite3, asyncio, email.parser, http.client, urllib.request, xml.etree.ElementTree, zipfile, tarfile, statistics and some popular third-party packages (numpy, pandas, requests, rich). The tool has three paths:
--version prints a string and needs nothing;mean reads a small CSV file with the csv module and calls statistics.fmean;stats loads the same file with pandas and prints a summary.The lazy version of the tool is identical, except that I prefix the imports with lazy. It is a one-word change per line.
I use Python 3.15 (3.15.0b4) on an Intel Xeon Gold 6548N (Emerald Rapids), with numpy 2.5, pandas 3.0 and requests 2.34.
| command | eager imports | lazy imports | speedup |
|---|---|---|---|
--version |
295 ms | 20 ms | 15× |
mean |
296 ms | 24 ms | 12× |
stats |
299 ms | 224 ms | 1.3× |

Printing the version number goes from 295 ms to 20 ms. If you subtract the interpreter startup (11.7 ms), the cost of the imports goes from 283 ms to about 8 ms: a 35-fold reduction. The mean command, which needs a few standard modules, is twelve times faster.
There is a downside: errors move. With an eager import, a missing module fails at startup. With a lazy import, it fails at the first use, perhaps deep inside a function, perhaps hours later in a long-running server.
2026-10-07 10:07:05

I write a lot about data parallelism. About how our processors have special instructions that can do several operations at once. They are often critical to get the best performance out of a CPU.
Compilers are sometimes able to take advantage of these instructions. But it is hard to tell when it will work.
For Java programmers, you can write directly for data parallelism since JDK 16. With enough skill, you might multiply the performance of a few key functions.
The Vector API is in the incubator module jdk.incubator.vector. It still not officially supported, but you can enable it by passing --add-modules=jdk.incubator.vector. I expect that it will soon become a mainstream feature in Java.
Roman Snytsar wrote a book about it: Mastering SIMD with Java Vector API, and the subtitle is Unlocking Single-Core Performance Through SIMD Optimization. I had the honor of being the technical reviewer. I learned a few things while reading the book and I enjoyed it.
Who has time for books, especially technical books?
What a book offers is time to reflect. Reading technical books today is probably just as relevant as it ever was. The author takes you on a story… and gets you to think.

Snytsar’s book works from problems like sum an array, compute a mean and a standard deviation, remove duplicates, merge two sorted arrays, and so forth. A lot of them are similar to the problems you will encounter as a programmer. He starts from the ordinary solution and rebuilds it with the Java Vector API. He then runs benchmarks. He looks at the assembly code.
The book covers performance issues such as unrolling, asymmetric loops, structural hazards, dependencies. Snytsar shows us that, in some instances, data parallelism can fail to bear fruits. The negative lessons are just as important as the positive ones.
The Java Vector API is a thick layer of abstractions, but Snytsar shows that some tricks work better on some hardware than others. I liked the chapter on removing duplicates, which is built on compress. It is one instance where the specific hardware matters. And you are unlikely to just find out about it yourself.
It is a book worth buying if you are a Java programmer with a focus on performance.
The book: Roman Snytsar, Mastering SIMD with Java Vector API, Apress, 2026.
2026-10-06 16:00:20

When you build a program, the compiler turns each source file into an object file. Then a linker stitches all the object files and libraries into one executable.
On Linux, the default linker is usually GNU ld (also called bfd). The GNU binutils also ship gold, an alternative linker that was designed to be faster. Gold was built a Google and first made available in 2008.
The mold linker is a newer linker written by Rui Ueyama, it was first released in 2021. It is meant as a drop-in replacement for GNU ld, and its main selling point is speed. The latest release is version 3.
How much faster is it on a large project? I built Node.js from its main branch on an Intel Xeon Gold 6548N server (Emerald Rapids, two sockets, 64 cores and 128 threads) running Linux with GCC 14.3. The final node executable weighs about 160 MB.
I captured the command that links the node executable and ran it six times with each linker. I report the median.
| linker | time to link node |
|---|---|
| GNU ld (bfd) 2.41 | 2.52 s |
| GNU gold 2.41 | 1.46 s |
| mold 3.0.0, 1 thread | 0.49 s |
| mold 3.0.0, 8 threads | 0.13 s |
| mold 3.0.0, 128 threads | 0.11 s |
With multithreading, mold is about 24 times faster than GNU ld and about 14 times faster than gold.
Even if I restrict mold to a single thread, it links Node.js in half a second, five times faster than GNU ld. With only 8 threads, you get nearly all of the benefit.
Using mold is easy. You can pass -fuse-ld=mold to GCC or clang. Or you can wrap your whole build:
mold -run make -j128
The -run option intercepts every call to the default linker and redirects it to mold. You do not need to change the build scripts.
Of course, two seconds saved on a link does not matter much if you build Node.js once. But a developer who edits a file and rebuilds dozens of times a day pays the link cost each time. With mold, the link step becomes nearly free.
My scripts and raw results are available.
2026-10-05 16:00:51

2026-10-05 01:07:27

Much of my research is on arXiv. I was one of the early adopters. It is simply a repository of research papers with PDF and metadata.
It makes it convenient to find research without any paywall.
Physicists started it in 1991. Math and computer science followed gradually.
It relies on volunteers to moderate it, because we don’t want garbage or spam.
On October 1, they capped the submissions to two per person per month.
Why?
arXiv got over 40,000 submissions in September.
Two years ago, it was about 20,000.
Doubling every two years. It is unsustainable for human beings.

In related news, Google stopped taking new bug reports in its open-source bounty program. They could not cope with the influx.
The writing is on the wall.
2026-10-05 00:49:29

Google’s first orbital data center is in orbit (October 1st). They say it is about the size of a fridge.
It has one kilowatt of solar power. It is about the power of a household.
It was launched on a Falcon 9 rocket (SpaceX).
It should run for about a year. Google plans to launch two more orbital data centres next year. The plan is to use laser communications.
For people complaining that the latency is going to be a problem with orbital data centres…
An orbital data centre can be effectively in your line of sight and it can be relatively close (much closer than the diameter of the USA).


Further if you have a cluster of space data centres, they can remain close and in a line of sight.
So you are not going to run your video games or do high frequency trading from space, but for compute, latency is unlikely to be the bottleneck.
The great features are space (no need to wait five years for a permit) and power (no need to negotiate with a local government for electricity). Solar power in space is great because it runs 24h a day.