MoreRSS

site iconKevin KellyModify

Senior Maverick at Wired, author of bestseller book, The Inevitable. Also Cool Tool maven, Recomendo chief, Asia-fan, and True Film buff.
Please copy the RSS to your reader, or quickly subscribe to:

Inoreader Feedly Follow Feedbin Local Reader

Rss preview of Blog of Kevin Kelly

Arguments in Favor of AI Fair Use

2026-09-07 19:00:00

I made some notes on the nature of the training of LLMs, and about whether we as a society should consider the material used to train them as a fair use of that material. I believe it would be best for us to consider them fair use, and I made some points in its favor. The context of these points is US copyright law, which I did not spell out here. There are other commonly used arguments pro, and many arguments against, which I have also not listed. I wrote these points as a way for me to think aloud along my own lines, and they are not arranged in any order to make a tight case. I welcome constructive comments, or points I may have missed.

  1. Text, images, or music, etc, are all expressions made by humans, which are copyrightable. During training LLM transforms these expressions into an abstraction called “latent space.” It takes one kind of thing – expressed content – and transforms into another kind of thing – an almost mathematical abstraction that contains no expression. This is why LLMs are called transformers. The latent spaces of an LLM are closer to something like a syntax which is not copyrightable.
  2. This latent space transformation is a fundamental transformation, because it goes one-way. It is an asymmetrical process: Content can be moved into latent space, but the latent space cannot be reversed to go back into the original content. It has truly been transformed.
  3. The usage patterns of how the general public uses the transformed material is new. When an LLM is asked to troubleshoot a broken electronic device, this service does not resemble reading a book or a newspaper article. When an LLM is used to develop a marketing strategy, it is a profoundly different action than reading text. When an LLM is assigned the task of proving a math problem, it is not substituting for a book about math. The transformations accomplished by LLMs are both in their form and in their use. Not only are they transformed, but they are transformed into a whole new category we have not made before.
  4. The key benefit of LLMs is their smartness, which emerges from the latent space and is not present in the source material itself. Their ability to pass a history exam is not present in any history book, and is not even present in ALL history books. These new benefits emerge from the process of creating an LLM trained on history books, and then transformed into a latent space, which produces this smartness. Much, if not most, of the value of LLMs derives from their ability to generate new benefits beyond what the source material alone can yield.
  5. The information inside an LLM is not stored in the form of copies of things. Even though an LLM may know the full content of millions of books, it does not contain within it any copies of the books. It may be able to recognize any object in a picture, even though it stores no pictures. Instead of containing a copy of the image, it contains the information contained in the image. The weights of an LLM can be copied, but the latent space itself does not resemble a copy. The latent space is a non-copy entity.
  6. Intermediate, transitional copies are the norm in the digital world. When your phone pulls up a web page it is technically making a copy of that page for a brief moment. When you send an email, it is copied in transit by telecom companies many times, but we don’t count them as invoking copyright. Since these intermediate copies are not stored, we don’t constrain them. Courts have already permitted literal copying that never surfaces publicly, The copies that LLM’s read once during training, are not stored, are not surfaced, and are therefore normal transitional copies that are fair use.
  7. The genius of the LLMs comes in part from the astounding fact that all bits of information from millions of different kinds of sources and subjects are mapped onto a single “map”. There is a single conceptual space (the latent space) where every fictional story, and every bit of medical information, is combined with all geological knowledge, and all news events. And so on. As a result the value of any particular contribution to the overall value of the LLM is almost incalculably small.
  8. Any individual item used in training the LLM has a very low value because 99% of almost every source is redundant. There is very little unique information in any item, even for seemingly original material. When an additional piece of content is inserted into the conceptual space of all human knowledge, most of it is redundant with other material. That means, removing an average bit of source content would make virtually no difference to the output of the LLM. As LLMs continue to scale up, the value of any particular source continues to diminish.
  9. Modern LLMs have trillions of parameters – that is they have trillions of attributes they are describing. The consequence of mapping all human knowledge into one latent space with trillions of attributes is that every bit of information intersects, or influences, every other bit of information. Calculating this enormously complex relationship between trillions of influences is the grand task of a massive data center. This calculation is the most complex computing task we have ever done. But it is so complex that we cannot practically unravel the ripple of influences of any particular source. The degree of influence for each source content is therefore not calculable at this scale. It is not dissimilar to our considerable inability to calculate in a quantifiable way the influences in our own thinking and knowledge. Therefore the import of a particular source cannot be practically measured. 
  10. In creating its knowledge, an LLM does not distinguish between sources of influence in its training. There are more hours of Star Wars reviews than Star Wars movies. You will have a lot of knowledge about Star Wars even if you have never watched a copy of a single one. What an LLM knows about Star Wars may come primarily from sources other than the movies themselves, from authors who have watched Star Wars. These diverse secondary and tertiary influences are doing a lot (most?) of the work that LLM delivers. Verbatim quotes can even come from secondary sources, which themselves are using the primary material in a fair use way. In fact the high value of LLMs derives in large part because indirect sources of material give a necessary and “intelligent” context to its knowledge. In this way, the LLMs improve the primary source material, to make it truly useful. This is another way LLMs transform material.
  11. When a human student is studying, they might learn from a book they purchased or from a book they borrowed from a library, which we consider fair use. We don’t judge their learning based on whether they paid for the source, or whether the sources were borrowed or rented. Likewise if an LLM is trained on borrowed copies from a library (fair use) that should not impact our judgement of its knowledge.
  12. The constitutional rationale for US copyright is incentivizing creation for public benefit, not necessarily securing revenue streams. I’ve personally seen anecdotal evidence that LLMs have been encouraging massive degrees of co-creation. Nearly every day someone sends me something that they have recently created with the aid of an LLM. You might dismiss this as slop. Most of it is. But most of what humans traditionally made everyday could also be called slop. And the best of the slop will get better. What is clear, is that the public benefit of transforming expressions into latent space is huge, and should be encouraged.
  13. At the moment the smartest AIs humans have invented are LLMs, but that might not be true in the future. There are lots of alternative models to LLMs being experimented with, and they could become the default models next. Today the most advanced LLM models are intensely commercial (and rich). But they might not always be. We already have various types of open weights, open source LLMs. We tend to think of LLMs as chiefly commercial entities, trying to maximize revenue. But LLMs don’t need to be profit-making. Many of the open source models are matching the capabilities of for-profit models today. LLMs can be non-profit. They can also be considered a true common wealth, like a public domain, or like the internet. I call one version of this kind of AI,  Public Intelligence. There could be more than one public intelligence. These LLMs would have multiple stakeholders, be publicly funded and publicly accountable. They could be trained on all texts in all languages from all times. In part they would be like a super library with a super librarian, able to do most knowledge work. This public intelligence would be available to anyone, at cost. The latent spaces of the LLMs (whether profit or non-profit) are our new public domain. The purpose of copyright is to maximize the public domain by encouraging abundance of creations with temporary monopolies and then moving the protected into the commons as soon as possible to assist in further creation. The new LLMs are a new kind of commons that provide society with new powers of understanding (getting fantastic answers to questions), new ways of self-improvement, new ways to collaborate, new ways to be creative. It would be best for society and everyone born and yet unborn, if public LLM commonwealths were encouraged. We should maintain the fair use option for training LLMs to ensure we have at least some public intelligence commons.

Weekly Links, 09/04/2026

2026-09-05 05:10:00

A Billion Dollars

2026-08-24 19:00:00

What would you do with a billion dollars?

I am occasionally asked by young people for my life advice. Could I counsel them as they decide what to do? They are thinking of going into something new. Or thinking about quitting their job. Should they work for a non-profit? There is usually a financial undercurrent to their concerns. I have a pretty standard response; I ask them to explore this scenario:

Imagine I am a wizard with a magic wand. I wave it over you and now you have a billion dollars. What would you do with your life? The price of things is no longer material. What do you use the money for?

The answers are wild and vary. They first buy a house, and maybe one for family members. Maybe even a second home at the lakeside. They get a car they always wanted. Travel some. Support an artist friend.

I tell them, that is all good, but after all that you will still have a billion dollars. You’ve only spent millions and you’ll make that back in a day or two. And you still have the problem of what you are going to do with your precious time.

Here the answers get interesting. Some would start their own business, or open a bakery. Work with doctors serving the poor. Go pro with their band. Become a YouTube creator. None of this requires anything close to a billion dollars.

I am reminded of a scene in the movie Wall Street where Bud, the protagonist, is working in a high-paying finance job he hates so that he can make a million dollars, retire and ride a motorcycle across China. The inside joke is that you don’t need to work on Wall Street, or to make a million dollars, to buy a motorcycle to cross China. You could easily do that whole trip for $2,000 dollars, saved from a short summer stint working at Costco.

The point of my exercise is that most people’s dreams are not gated by money. Sure you need some money, but you probably don’t need as much as you might imagine. You most likely don’t need multi-million dollars to start your dream.

The bottleneck for most dreams is not financial. It’s a lack of confidence, willingness to take a risk, and face failure, or a lack of imagination of what is possible with current resources. If you want to write a novel you don’t need capital — you need discipline and 1,000

words a day. If you want to start a consultancy you need expertise and a first client, not a war chest. A lack of money is a convenient thing to blame because it feels concrete, but it’s rarely the actual blocker for achievements.

I am not suggesting money is not necessary, or that you should go into debt to pursue your desires. This exercise assumes you’re not choosing between your dream and your next meal. If you are, I would use my wizard’s wand to first grant you a safety net — everything after that is what this essay is about. Money is the fuel you need for a trip – but it is not the goal of the trip. You most likely don’t need more money to accomplish your dream. You need time, skill, relationships, courage, and perseverance.

I’ve managed to do a lot of things without having much money at the start. Here are some tips that I have used:

Start Small, Scale Later – Whatever the cost of a dream, you usually don’t need it all up front. You can begin with a little and if the path is a good one for you, you can gain the additional resources as you proceed.

Constraints Breed Ingenuity — Forward motion and ingenuity often open up new possibilities that can route around costs. This is why most breakthroughs happen in startups. If you could ensure a breakthrough by spending money, all the new inventions would only happen in big corporations. They don’t, because you can’t buy breakthroughs. Startups have no money to buy a solution, so they are forced to be ingenious, clever, thrifty and bold to invent one. They have no choice but to spend their imagination and creativity. That outside thinking generates the new ideas. It is often the financial constraints that make dreams come true.

Small Tweaks, Big Difference – If you are an average worker, you’ll channel more money in your lifetime than you think. The typical median American will see several million dollars of income flow through their accounts over their life. Small tweaks in handling those quantities can make a big difference in achieving your dreams.

Prototype Your Life – Try a scaled-down version of your dream first. Start with a pop-up instead of a storefront; or a self-published ebook instead of a commercial press deal; some hand-made versions instead of a factory run; volunteer to intern with a lawyer before heading to law school; do a self-contained weekend camping trip before tackling the Appalachian Trail. The prototype will reveal what your actual monetary costs might be, and what the real bottlenecks are. Prototyping will often adjust your dream and save you a fortune.

Relationships Beat Capital – Friends and relationships are more important than money. If you want to go sailing, a friend with a boat is better than you owning a boat. Investing into relationships is a better deal than investing into crypto. The most valuable things in life – the things your dreams should include – are actually things that no amount of money — especially a billion dollars – can buy.

1,000 True Fans – My theory of a 1,000 True Fans is aligned with this perspective. It says you do not need a billion fans, or a best seller, or a billion dollars, to succeed in your creative endeavors. To make a living you only need thousands of super fans, if you have direct relationships with them. The key is the direct relationships with your fans.

Lower Barriers – The “minimum viable capital” for most creative or entrepreneurial dreams has collapsed over the last decade. Because of new technologies, the barriers to making stuff have lowered. It is cheap and easy to print a book, to 3D print an ingenious device, to generate a film. You simply need a whole lot less money to make something compared to previous generations.

Other People’s Money – There may be dreams where a lot of money is required at some point: maybe you want to do a hardware chip startup. This is where relationships are important again, because you want to use other people’s money. And the big secret in raising other people’s money is that it is ALL about relationships and character.

Take a small amount of money, and leverage it with great ambition, imagination, and creativity. Try something that no one has tried before, where the solutions are not for sale. Make something that no one is selling. Become the world’s expert in something – anything – and use that expertise as a platform to do more things you like to do. Cultivate friends and make their relationships your capital. Focus on being valuable, helpful, interesting, weird, rather than on being rich. I love what Brian Eno said, “If all I’d ever wanted to do was make money, I’d probably be really poor by now.”

I’ve had the privilege of hanging around a dozen actual billionaires. Despite the myth, they are not unhappy. But here’s the thing: they are still trying to figure out what they want to do next, and who they want to become. And their billions of dollars don’t help them in that. In fact, their billions of dollars are almost a prison. Their previous success hemmed them in, set unhealthy expectations, and will never go away (it is hard to get rid of a billion dollars). The hurdles they have in achieving their new dreams are exactly the same as yours: the need for will power, perseverance, ingenuity and imagination. None of these cost a billion dollars; all of them are available to anyone anywhere.

The goal is not to acquire, but to become. At your funeral people will remember what kind of person you were, not what kind of stuff you had. Your everyday life will need money, and no matter what you do, a fair measure of it will flow through your days. But beyond some minimal amount (way less than a billion) money is not important to who you become. Other assets like your character, your ethics, your work, your relationships, your word, your spirit, your drive – these are your real treasures. Use them to achieve your dreams.

The wizard’s wand is a magic trick. You already have everything the wand would have given you — everything except the billion dollars — which, it turns out, you didn’t need. Nor do you need a magic wand to materialize these powerful assets. They are in your hand right now.

Without a theory of intelligence

2026-08-17 19:00:00

The history of science — and of progress — is a series of benefits that are direct results of new tools. A large part of our current longevity is due to the invention of the microscope. This new way of seeing opened up the microscopic world, a teeming universe we had no idea about, and from that view quickly came germ theory, and soon after, new ways to avoid many common fatal diseases. Hundreds of other advances were also birthed by the microscope, including our understanding of DNA. Similarly, the telescope not only opened the heavens to our inspection, it had a direct role in helping us devise the laws of physics, which in turn permitted harnessing the atom, developing GPS, and making cheap computers. Oscilloscopes, volt meters, barometers, cyclotrons — these are more than measuring tools; they are portals that open up new territories to be explored.

We are on the cusp of inventing a new cyclotron: artificial intelligence. Of course AI will usher in new ways to do stuff. We can offload chores we don’t want to do, but the greatest power will be in accomplishing things we had never imagined doing before. That new superpower will gradually revamp our society as we learn how best to employ it.

But a secondary revolution will come from using AI as a microscope: we will use it to see our own minds. AI will be a cyclotron that lets us inspect the mysterious particles of cognition swirling in human brains — bits that are completely invisible to us now. With this cyclotron we will be able to dissect thinking, take intelligence apart to see its components. We will be able to run endless experiments on whole and partial minds, experiments we can’t and won’t run on ourselves.

The space of possible minds in the universe is vast. Using AI as our scope, we will begin to populate that space, constructing artificial minds for specific purposes — ones that do math proofs, others for everyday robots, others to write stories, another to manage the planet’s climate. With the ability to look into minds, we can begin to build a theory of mind. Centuries ago the microscope gave us germ theory and cell theory. The telescope gave us gravity and relativity. The new scope of AI will give us a theory of intelligence.

Despite the great strides scientists have made in producing AI, we have no theory of intelligence. We don’t know how intelligence works, in humans or in machines. We don’t know what it is, or why it produces smartness. What is the smallest possible thing that will produce intelligence? We don’t know. Is there a universal ingredient shared between humans and machines? We don’t know. What is the metric for intelligence — how would we even quantify it? We don’t know. If we want more of it, what’s the formula? We don’t know. Is intelligence one thing or several — one dimension or many? We don’t know that either. Until we can predict what AI will do, we have no theory, and if we have no theory, we can never predict what AI will do. This is a problem.

Our ignorance about intelligence is vast. Take energy: does intelligence require a lot of it, or just a little? The supercomputer inside our own skull runs on about 25 watts, which hints that the minimum energy needed for intelligence might be small — and that the giant, controversial AI compute centers we’re building now may turn out to be a temporary blip rather than a permanent feature.

In the 1800s, the field of artificial power — steam engines — was hobbled for decades by the lack of a theory of heat. Tinkering engineers spent years on inefficient machines, costly detours, before arriving at the insight that power comes from temperature differential, not temperature alone. There were still engineering problems to solve after that — inventing high-pressure equipment, for one — but the theory told them which hardware problems were worth solving.

Like the theory of heat, a theory of intelligence will require engineers and theorists working side by side. It will be a nerd endeavor that requires building things in order to understand them. It will need a diverse team — neurobiologists, physicists, mathematicians, computer scientists, philosophers, hardware experts — and I expect a lot of foolish ideas will be pursued in order to discover general principles of intelligence.. This is basic science at its purest. It will take patience.

The payoff is a much better sense of where to focus attention and resources. We would know in advance, not after a $100 million training run, whether a given approach is near-optimal, or where the cognitive waste is happening. A theory tells you where the limits are. For instance, it could answer this: are today’s LLMs 1 percent of the way to what’s physically possible with a GPU, or 90 percent? That would be extremely useful to know. And since we’re about to invent thousands of new kinds of minds, a good theory would let us reliably engineer the right configuration for a given task — the way thermodynamics informs the design of an engine.

There’s a good chance a theory of intelligence will turn out to be necessary for real alignment. We may need to understand what the internal representation of thinking should look like before we can engineer core values — not just rules — into an AI mind. Right now we’re doing this by unscientific trial and error. We are trying anything we think of to get AIs aligned with our tricky goals: be creative, but don’t do anything stupid or bad. An explicit formula of intelligence, with a sense of its tradeoffs and limits, would give us something to steer by.

But a theory of intelligence isn’t just for AI. We are thinking machines too, and our intelligence sometimes needs fixing. Modern medicine still classifies psychiatric and neurodegenerative conditions largely by symptom cluster (the DSM approach) rather than by which specific computational function has failed. Unlike brains, artificial systems can be opened, probed, disturbed, and modified to isolate function directly, at a level of access neuroscience has never had. I’d bet we learn more about our own brains from building a thousand AIs under a real theory of intelligence than we’ve learned from a century of neuroscience.

A theory of intelligence might also finally decouple intelligence from consciousness. Intelligence is likely a computational quantity; consciousness is a phenomenal one. Of course, there is a chance that we might discover there is no general theory of intelligence at all; that like life, intelligence is a squishy, messy, almost illusionary phenomenon that cannot be encapsulated into a mathematical formula. That would be unfortunate, but good to know sooner rather than later.

Right now a small group of scientists are trying to find a new field: the science of intelligence. Jacob Yates, a neuroscientist at UC Berkeley and one of the effort’s conveners, says there are already glimmers of where a theory might emerge — pointing to work connecting stochastic thermodynamics with variational inference, and to geometric tools for characterizing the promises of what can be learned from data. A theory might suggest that every unit of learning would need X amount of energy, and Y bits of data. It might propose diminishing returns on scale, or the theoretical limits on how smart anything can get.

Information theory transformed electronics, guiding and accelerating everything that followed. A theory of intelligence would do the same for work, research and science itself. However with or without a theory of intelligence, artificial intelligence is becoming our new cyclotron. It is a telescope that is opening up a new territory: the continent of minds. Once the most mysterious force in our lives – our minds, all minds – will then be revealed. Indeed, the most complex things in the known universe are now available for exploration and study. The entire realm of learning, smartness, and thinking will be near, accessible in the most practical way. In the long term the instrument of AI will probably exceed the importance of the microscope, telescope and cyclotron combined. We find it hard to see the real world without views of the ultra tiny and ultra large made possible by our tools. Future generations will find it hard to see the real world without views of all the possible minds operating upon it, each seeing the world in a slightly different way. A world without the tools of AIs will be unthinkable.

Weekly Links, 08/14/2026

2026-08-15 05:22:00

Worldbuilding with Spatial Intelligence

2026-08-10 19:00:00

The next stage in the development of AIs is to give them spatial intelligence.

Our current, smartest AIs are masters of words. They have been trained on zillions of words. Their education consists of the knowledge we have written down into new books and journals. They are smarty pants, the classroom genius that has read everything. But not only have the AIs read most books, they actually remember everything they have read. The LLMs today have a PhD level of knowledge in literally every subject, which makes them powerfully book smart.

But they often lack common sense, and when they are given a body, as in a robot, they flail, flounder, and stall because they have no embodied intelligence. They don’t know about reality. Operating in the real world takes a different kind of intelligence than book smartness. The smartness of a body as it moves in the world requires an intuition about gravity, and lightning fast visual perception, and an awareness of three dimensions – up, down, front and back – and many other basic responses that we humans have learned over millions of years of evolution and is baked deep into our reflexes. That kind of embodied spatial intelligence has yet to be trained into our AIs.

Many labs are trying to give AIs this missing spatial intelligence. Every major robot company is working on some version of this research. The prize for succeeding in equipping a robot with an embodied intelligence is monumental; we would finally have robots in our homes, offices, factories, and everywhere. They could get around as well as we can, fold a t-shirt, cook a burger, bathe an invalid. The AIs would do to physical tasks what they have done for intellectual tasks.

But in addition to unleashing the robot world, spatial intelligence would also unleash something else: worldbuilding.

World Models

Spatial intelligence would give us world models: AIs that have real world knowledge, not just book knowledge. Instead of being trained on words and descriptions of reality, as they are now, they would be trained on reality directly. They would witness the bounce of a ball instead of a description of a ball bouncing.

The primary bottleneck restraining the arrival of this world model is the lack of sufficient quantity of quality data. Large Language Models (LLMs) worked because the internet had already digitized language: libraries of books, all journals and newspapers, and years of public conversations and personal blogs. Petabytes of digitized text existed and were vacuumed up as training material for this model based on language. There are no equivalent sources of digitized petabytes of data derived from reality. If you are trying to model the world you need tons and tons of data about water moving in all its ways from waves, to splashes, to sprays, to streams and drips. You need data about clouds and wood as building material, and the way clay squishes, and traffic moves, and fabric falls, and balls bounce. You need real data from every corner of life, just as we have text about every corner of life.

The nearest deposit we have of this kind of reality data is the video on YouTube. YouTube has never disclosed how many hours of video they have but it is widely estimated to be in the billions. Given the approximately one million hours of video uploaded every day, the range of human activities YouTube captures is fairly large. There are endless hours of sports play, cooking in kitchens, people working at physical jobs, moments of everyday life, including millions of hours of accidents and improbable events, which are even more valuable when training a model. In addition, there are billions of hours of CCTV security camera footage, and the video recordings from car cams on the roads. These are also being used to train world models.

Many startups are racing to create foundational world models, but I think the first ones to succeed will likely be the platforms in control of this data bank. In the US, Google (owner of YouTube), and in China, ByteDance (owner of Douyin/TikTok), or those working in partnership with them. Importantly, the best source for a robotic spatial intelligence data will be the continuous experience of robots themselves. As robots work, they rapidly accumulate very good data about the real world that they are scanning. The more robots that have been turned on, in more locations and occupations, the more and better data they collect. Even limited, lame, or poor robots can collect good data, which gives great incentive to get robots out in the world. It will probably be economically shrewd to lose money on early robot models in anticipation that the data they gather will be worth more later in improving newer versions. This incentive might be so strong that robots are sold below cost to you as long as you keep using them.

AR and XR

The second significant source of world modeling data will come from smart glasses. Smart glasses have transparent screens you look through as well as cameras that look out. Wearing them you basically see the world that a robot sees. The cameras in the glasses scan the world ahead of you, helping the chips inside to render a digital version of the scene ahead which is laid over the real scene, so that you see a merged version of both. This enables software to render smart annotations to the real scene, whether they are navigation aids (follow the blue arrows on the ground), or fictional fantasies (follow the blue fox running in front of you). The result of melding AI generated scenes with real scenes is known as Augmented Reality (AR), or Mixed Reality (XR).

The crucial step for AI is that the cameras in the glasses in the millions are constantly scanning and re-scanning the world, feeding huge amounts of data about the real world into AIs to be digested and processed. Huge amounts of 3D AI is needed to perceive and “understand” the layout of the world, to recognize where something is, or to recognize what it is. All the deep situational awareness we expect from a pair of smart glasses requires world modeling. We can’t have AR, or XR without cheap, ubiquitous AI. Conversely, there probably is no 3D AI without cheap, ubiquitous smart glasses to train the models on. We might also expect some AI companies to sell smart glasses at a loss because the data they are generating might be more valuable than the cost of the hardware.

Three Stages of Digitization

I like to divide the digital world into three stages. In the first stage, we digitized information and made it machine readable. Machines could read, recall, search, and share all the information of the world. That happened during the dotcom era, and the owners of the machines became the dominant cultural gatekeeper. In the second stage, we digitized the relationships between humans. Machines could see, recall, search and process who was friends with whom, who dated whom, who worked for whom, who liked, who disliked, who swiped, who watched, who voted up – the entire realm of social relations was now machine readable. And the gatekeeper of those machines of the social networks dominated. We are now about to enter the third stage, where the entire physical world is machine readable. Once millions of people start wearing smart glasses, scanning, perceiving, ingesting, the world 24 hours non-stop, every bit, every action, all physical phenomena are digitized and available to be read by a machine (AI). That would enable us to process, search, manipulate the real world in the same way we did with information.

At the most trivial level we could search the entire world for example; find me a park bench with an unobstructed view to the west, where the sunset light in mid-December glints off the windows of a tall building in the shape of a teardrop, and it would search the known world for that configuration. But it could also search the world for visible evidence (vibrations, rust stains, cracks) that a bridge is in danger of failing. This real-time constant scan of the world, combined with an AI world model, generates a digital twin of the world. We can probe it, search it, for patterns, but we can also run what-if simulations and scenarios in it. This is the Mirrorworld.

This full-strength digital twin applies not just to the whole world but to all its parts. If enough people scan a building with their glasses as they work in it, and the workers maintaining it physically scan its behavior, these all together create a digital twin of that building. That 3D digital twin then becomes a tool for managing the building. The digital twin of the building – its mirrorworld version rendered by spatial AI – can generate predictions of what the building might do next, scenarios of what could be done with it as is, and reminders of what needs to be done next to keep it going. Plus the digital twin serves up an augmented reality for any visitor to the building.

The world AI model in smart glasses does three things at once. They display virtual worlds and virtual annotations. You put them on, and virtual smart things are added to what you see. So the generative powers of AI will create these visuals in the same way they generate fictional video clips. The 3D models can generate arrows, advertisements, virtual characters, avatars, text annotations, anything. At the same time, the second task the same AI world model performs in the glasses is to scan the real world in order to cast the virtual layer exactly. The generated images in AR and XR will match the lighting of the scene you are in, and the virtual annotations will geometrically fit into the scene perfectly. If your friend is going to appear as a 3D avatar sitting in the chair next to you, the AI needs to construct this synthetic melding of both worlds with a precise perception of the room you are in. So the cameras in the glasses are constantly scanning the world in order to make the display of the virtual parts believable. Thirdly, the cameras scan the world in order to gain more, and more up-to-date, data to improve the model’s intelligence. They scan not just to render a synthesized view, but to keep getting smarter. In this way, augmented reality (the mirrorworld) will train the AI world models.

The Internet of Things

The Internet of Things was a vision promised by digitization. The idea is that eventually every object, every artifact, every thing, is added to the internet. Each item in your home, your toaster and your washing machine, even your shoes, are connected to the internet, and become smart. All the objects in an office and factory are added. Every item on the shelf of a store is added, with their prices dynamically changing as they age, or are in demand. To achieve this internet of things, we would have to add a tiny chip inside each artifact produced, and maybe add power. Everything would get its own IP address. This did not happen, as plausible as it sounds.

However, spatial AI can create a version of the Internet of Things using the data from the mirrorworld of AR or XR. Smart glass scans everything in their view, so in the goodness of time, the entire world will be scanned and digitized. The outside of buildings, the inside of buildings; public spaces, your bedroom; monuments, and the things in your refrigerator. The nature of spatial intelligence is that it perceives the whole view; it can semantically structure what it sees, distinguishing a person here, a car there, a hat on a person’s head, a window in a building, the sill of that window, one pane of glass in that window, a sticker on that pane of glass. The AI can extract all of these patterns from the mess of reality. In other words, the AI is capable of identifying every object and part of every object in the world. It identifies singular things not by giving it a unique number (like an IP address) but in relation to other things. So that window sill is the sill in the window that is 5 up and 4 windows across from the main door of the building that is next to the post office in the city along the harbor at the mouth of the river, etc. That window sill, or that door, or that chair is connected to everything else by some semantic relationship that can be extracted by the AI. To the extent that we can reach the AI via the internet, which is always on, everything the AI sees is therefore on the internet. When systems are scanned and included in the mirrorworld, they also appear in a virtual internet of things.

The virtual Internet of Things – a system that knows about all objects and places – is a powerful tool. You could use it to simulate a factory, a city, a company, a household, or a store.

As chips become free and batteries ever better, many, many physical objects really will get their own IP address as well. Those chips will transmit internal states of things that scanning won’t reveal. But even these animated objects will be fitted into the model of the world that is created by constant scanning. A fully developed world model will be able to semantically parse the world, and just as an AI knows every page of every book, the spatial AI will know every room of every building, and every object in every room.

Similarly, this same spatial intelligence will be able to semantically parse the entire visual world we have recorded in video and movies. It will be able to superhumanly comprehend, remember, grok, search, find, and reconstruct the content in every frame in every video ever made. Find me all the moments when a white rabbit is pulled out of a magician’s hat. Where do tweezers appear as a close up in movie history? How has the shape of bathrooms changed over time?

This semantic knowledge of the world, both real and fictional, will unleash thousands of new services and products we have not imagined yet.

Worldbuilding

For instance, simple worldbuilding will become one of the new superpowers we get from spatial intelligence. Once the world’s detail is machine-readable the way its language already is, a director can generate a coherent world instead of filming one. Solo individuals will be able to create a feature-length film in their bedrooms. In the same way as a young talented J.K. Rowling could singlehandedly create a deep satisfying wizarding world in immense detail, young directors will create movies and games with deep satisfying details and drama with little additional help. Of course most of these will be unwatchable with an Audience of One, but this is true of text creations as well. Most novels, fantasy stories, or sci-fi, and books in general are not worth reading; most AI co-generated movies and games won’t be worth it either. But occasionally, one will be brilliant. And it will be marvelous and something that could not have been created via the old-fashioned way.

Worldbuilding is a major part of what the best fiction does. The author creates a world and brings you into it. In the past, to do it well was such a demanding task that few individuals could excel at it; it was often a group collaboration. Science fiction movies and games harnessed the best worldbuilding at great cost. But as it becomes easier and automated, worldbuilding should become a common endeavor. A what-if question could be answered by building a world. What if our product was in every car? What if everyone were given some stock at birth? What if we put cameras everywhere that anyone could watch? And then you build out an entire world to see what happened.

Worldbuilding is a type of simulation; but instead of trying to ensure the dynamics adhere to reality, you can feed it alternative rules. Try stuff. This fiction can be entertaining, but also useful. Worldbuilding assisted by AI could easily become a standard way of designing and managing complex projects. You build out an entire world based on your idea, and then immerse yourself in it to evaluate it.

The Trust Problem

The greatest friction slowing down the arrival of this mirrorworld, worldbuilding, and spatial AI is the tension that constant scanning and privacy surveillance brings to this tech. Already, many people are upset by the cameras embedded in current smart glasses. Terms like glasshole are applied to those who leave the cameras on in public. The degree of trust required to allow a company to record not only everything we see, but to let others record everything we do in public is beyond the limit for most people today. It seems unlikely we’ll change our attitudes about being constantly scanned any time soon.

I think we will change our attitude, gradually. The main driver will be the benefits we get. I bet we eventually will become oblivious to being recorded in public. Residential and commercial buildings have cameras around the property, cities film public spaces, and cars – especially self-driving cars – film everything around them. Having cameras on people’s glasses will not feel so out of place. Additionally, the kind of scanning done for AI might come to be seen as very distinct from a traditional recording. In fact, the streams of images that are used for training will probably not be saved, in the same way that the text for training AIs is not saved once ingested. The smart glasses may be on, capturing images, but not saving the stream for later reviewing. Furthermore it may be recognized that having an AI watch everything is different from humans watching. Billions of people have become comfortable in having AIs read all their email on gmail; it is not the same as humans reading all your email, and there are benefits to Google “reading” it all. There may be ways to anonymize our own captured behavior, so that we can benefit from its digitization, but not be concerned that our privacy is compromised. Technically this is possible; but it requires trust in corporations to execute this process reliably.

AIs in general have trust issues, so the advent of spatial intelligence will depend on how we resolve our trust in large corporations. AIs will remain black boxes, hard to understand and hard to predict. Spatial AI, smart glasses, and robots will share some of this uncertainty, as well as the challenges of maintaining a sense of privacy.

However, the benefits of spatial AI and world models will be huge, and that helps us overcome our fears. From these mysterious models will come real working robots in the millions, engines of wow generating movies and games by solo individuals, a new social media of convincing avatars in immersive 3D presence, augmented mixed reality, and a thousand other things that exceed my meager imagination.