2026-08-29 00:45:11
Recursive self-improvement, where AI continuously builds better versions of itself, might be harder than some hope.
There’s growing excitement in the AI industry about the idea that today’s leading models could build the next generation of the technology. But a new study recently found top AI agents struggle on the kind of genuinely open-ended research problems required to push the field forward.
Large language models have made rapid progress in many of the day-to-day jobs involved in machine learning research, such as writing code, generating and curating data, and running experiments. Last year, startup Sakana AI’s AI Scientist-v2 even managed to write a paper that cleared peer review for the prestigious International Conference on Learning Representations.
These advances have led to speculation that models are close to being able to build better versions of themselves with little human oversight—a process called recursive self-improvement. The idea underpins predictions that we may be on the verge of an intelligence explosion that could quickly lead to AI superintelligence.
In a recent paper, researchers put the idea to the test using a new approach they call shadow evaluations. This involves taking the research question from a high-quality, unpublished machine learning paper and asking AI agents to solve the problem. The original paper’s authors then grade the results. When the team tested Claude Opus 4.8 on two papers submitted to the prestigious machine-learning conference NeurIPS 2026, the authors rejected both.
“The papers were nowhere close to the mark when it came to being at the quality of a top AI conference,” Sayash Kapoor from Princton University, who co-led the study, told MIT Technology Review.
Previous efforts to get AI agents to do machine learning research have often targeted problems focused on engineering, such as reproducing previous research or training smaller models against a benchmark.
In the new experiments, the researchers challenged models with more open-ended tasks that required them to devise hypotheses, decide what evidence is needed to validate them, judge when a research direction was fruitless, and go back to the drawing board.
One research question was whether the personality traits a language model displays can be measured and adjusted by observing and editing its weights; the other attempted to detect when a model that works with tabular data has quietly stopped being reliable.
In each case, the AI researchers were given $3,000 of API credits, a budget for time on GPUs to run machine learning experiments, a dedicated Linux virtual machine, and unrestricted internet access. They were then given six days to produce a paper that could pass NeurIPS’ stringent peer-review criteria.
In both cases, the models got a good start. The agents surveyed the literature effectively, came up with opening hypotheses that mirrored those of the authors, and successfully ran hundreds of experiments.
But they quickly went off the rails. Although they could monitor their own use of time and their API and GPU budgets, they rushed through the process. One left 110 hours of unused time on the clock, and both failed to spend even 50 percent of their API budget.
Both agents also settled on a research direction within just 10 hours and failed to change approaches despite repeated negative feedback from another AI designed to review drafts of their papers. The reviewer identified problems the human authors would also flag in the final paper, but the models simply added caveats to their findings and ploughed on. Ultimately the papers received a “strong reject” and a “reject” decision from the human reviewers based on NeurIPS grading protocol.
The authors admit their approach has limitations. The reviewers knew AI had written the submissions, and some of the team are on record as doubting an imminent intelligence explosion. The original human-authored papers also took far longer than six days to produce and used many more GPU hours to reach their conclusions (though, as the researchers note, the models did not use their allocated budget in any case).
Nonetheless, the results suggest that today’s models still have some way to go before they can tackle the most challenging problems in machine learning research. Until that happens, the dream of recursive self-improvement is likely to remain a distant prospect.
The post Are We on the Verge of an Intelligence Explosion? Maybe Not. appeared first on SingularityHub.
2026-08-28 05:13:16
Can AI replace lawyers—at least in some circumstances?
Last week, Australia’s Fair Work Commission ruled Gregory Baker, a computing academic at Macquarie University, should be treated as an ongoing, part-time employee, after the university had earlier declined his request to convert from a casual role.
It was immediately described as a “landmark” decision, the first test of Labor’s “employee choice pathway” reforms passed in 2024.
But the ruling also made headlines for other reasons. Baker represented himself at the tribunal and has said he won with the help of trained artificial intelligence agents. His success again has us asking: Can AI replace lawyers?
On closer scrutiny, Baker’s case looks less like evidence of AI replacing lawyers and more like a powerful illustration of how a highly capable user can employ AI tools to terrific effect.
Speaking to the Australian Financial Review following the ruling, Baker said it was actually an AI tool that alerted him to the possibility of converting his role from casual to permanent part-time in the first place.
He had been teaching computer science at Macquarie University over consecutive semesters from 2023 to 2025, and in November 2025, he gave the university the required notice that he believed his work no longer met the requirements of casual employment.
The university did not accept this notification and Baker lodged a dispute at the Fair Work Commission—without a lawyer—in December 2025. The parties could not reach agreement, and the case went to arbitration on May 12. A decision was handed down last Wednesday.
Baker has said he won by using multiple paid AI agents, such as OpenAI’s paid offering, ChatGPT Pro. This “team” helped assemble his case, follow up references, and anticipate his employer’s counterarguments.
His victory has been celebrated as historic, with the Australian Financial Review describing it as “the first known successful use of technology by a self-represented person in the legal arena.”
However, a few things set this particular case apart. Baker’s IT background, expertise managing AI agents, and ability to optimize their use for his case represent a rare level of expertise in using AI in a legal context.
Details included in the Fair Work Commission’s decision also suggest he kept his legal argument narrowly focused on teaching he’d done in one particular unit.
Less expert use of AI in court often sees those bringing claims produce “kitchen sink”-style arguments, which include weak, exaggerated, and nonsense claims.
Baker’s dispute was also narrow, limited to the application of a casual conversion law that had not yet been tested. Importantly, the Fair Work Commission (a tribunal, not a court) is designed to be user-friendly, to enable workers to bring claims without a lawyer.
Elsewhere, the use of generative AI in legal proceedings is attracting a lot of attention for less positive reasons.
Most of this attention centers on the damage caused by inaccuracies, hallucinations and “AI slop”, and how courts and tribunals should best respond.
By making it easier to put a case together, AI has removed traditional access barriers for some litigants. But while case numbers are going up, case precision and quality is going down, making it harder to manage disputes to resolution.
Courts and tribunals are struggling with the volume. At the Fair Work Commission alone, workload has reportedly increased by 70 percent over three years.
New challenges are emerging as time goes on. Reports suggest litigants and lawyers in some overseas jurisdictions are embedding prompts in digital documents (something called “prompt injection”) to overcome or manipulate AI-based review systems some courts use to process documents.
Baker’s example shows us something significant. Used well, AI tools can empower people with narrow legal disputes and digital skills to achieve successful resolutions at low cost.
This is an important development in access to justice. In Australia, there is a huge gap between the number of people with legal problems and the very limited funding available for legal assistance.
Most of the community is in the “missing middle,” unable to afford private legal assistance but on incomes too high to qualify for free legal aid.
AI tools stand a good chance of helping people with sufficient legal capability with problems and cases—like Gregory Baker’s—that are a good fit for the solutions AI can offer. These are few and far between, however.
We should expect case numbers and self-representation in courts and tribunals will continue to grow and expand beyond Fair Work.
While there will be some baseless cases, the growth also represents the natural consequence of removing one traditional access barrier to our formal justice institutions—getting in the front door to start proceedings.
The bigger picture challenges are persistent and raise important questions. Who will most benefit from the capacity of AI tools to enhance access to justice, and who will continue to struggle to get basic legal problems resolved?
For courts and tribunals, the challenge will be striking a balance between managing caseloads and delivering justice, while not wasting the opportunity to expand access to justice.![]()
This article is republished from The Conversation under a Creative Commons license. Read the original article.
The post An ‘AI Legal Team’ Has Won Its First Case. It’s a Rare Victory for Access to Justice. appeared first on SingularityHub.
2026-08-26 04:35:27
These lab-grown balls of brain tissue could help researchers study a host of disorders that emerge as the brain ages.
Five years is an eternity for brain organoids. Also called mini brains, these blobs of tissue have taken neuroscience by storm for their ability to capture the intricacies of developing brains.
Organoids begin life as a collection of stem cells. Within weeks, they spontaneously produce a range of brain cells. Neurons form circuits that spark with electrical activity. Gene expression resembles that of early fetal brains. Some organoids learn to control small, isolated muscles. Others link to spinal cord organoids and process pain signals.
Over time, they grow more sophisticated in both structure and function—eerily similar to near-term fetuses—prompting bioethicists to ask if they could one day become conscious.
But time isn’t on their side. Most mini brains survive only a few months before their sensitive neurons start to wither. Circuits break down, structures collapse, and eventually the organoids die. As a result, they can model only the early stages of human brain development, leaving what happens during the later months of pregnancy and after birth largely mysterious.
These periods are especially relevant to schizophrenia, epilepsy, severe autism, and a host of other disorders. Scientists have studied late-stage development using donated tissue, but samples are scarce and raise ethical concerns.
A team led by Harvard’s Paola Arlotta is now pushing the boundaries with organoids. Last week, they described a method that kept mini brains alive for over five years—the longest yet—and tracked their development throughout. Despite growing outside the body, the organoids matured on a timetable similar to normal brains. Genetic activity in the oldest ones resembled that of a typical 4-year-old.
The findings were originally reported in a preprint and have now been peer-reviewed and published in Nature.
The developmental lockstep surprised the team. Cells from older organoids, when mixed with younger ones, continued maturing on schedule, suggesting they carried an internal developmental clock that keeps track of their progress.
“The brain doesn’t develop in a vacuum. It’s an organ of incredible complexity that interacts with so many other systems,” study author Irene Faravelli said in a press release. “It was not a given at all that our simplified model would match natural development in this many ways.”
Because mini brains generate nearly the full range of human brain cells, they’re promising models for the study of early brain development. But early versions survived only a few weeks. Without blood supply, cells at their centers starved and died.
Through trial and error, researchers learned to coax them into increasingly sophisticated structures that included layers resembling the cortex and had integrated blood vessels. This vastly extended their lifespan.
In 2021, a study kept mini brains alive for up to two years, capturing cortical development from pregnancy to roughly a year after birth. Four years later, Arlotta’s team announced a way to extend organoid lives to a staggering seven years. Roughly the size of a pea, each nugget was packed with some two million healthy neurons and other brain cells.
Following these organoids for years offers an unprecedented window into how the brain grows and wires itself—and how genetic changes early on might contribute to diseases later in life.
Our brains take roughly two decades to mature. Throughout this period, neurons constantly rewire their connections. Scientists have long known that conditions such as schizophrenia and some forms of epilepsy first emerge during adolescence. Because mini brains can be grown from a person’s skin cells and retain genetic mutations associated with neurodevelopmental disorders, they offer a way to probe how, and when, neural wiring goes awry.
But timing matters. The question is, how faithfully does a growing blob in a dish follow the developmental journey of a human brain?
To answer that question, the team grew 34 organoids and tracked them at regular intervals. They collected data every three to six months for the first 18 months, then annually until the organoids were over five years old.
Crucial to the brain blobs’ longevity was switching the growth medium—a nutrient- and protein-rich slurry—halfway through development. The new recipe kept neurons alive longer, giving them time to support increasingly complex activity.
The team then tracked changes in gene activity and epigenetic markers (chemical tags that control which genes are turned on or off). They then compared the findings with data from younger organoids—ranging from 15 days to six months old—and donated human tissue.
The developmental timeline was surprisingly similar to that of a human brain. Young organoids showed gene activity resembling the first trimester; by three to six months, they looked more like second-trimester brains. After a year, their gene activity profiles resembled those of newborns. By the end of the experiment, they most closely matched a typical 4-year-old.
The team also tested them with epigenetic methods used to gauge biological age as opposed to calendar years. The organoids gained and shed epigenetic markers in patterns that broadly tracked those seen in natural brain development.
The organoids seemed to retain a “sense” of time. The team mixed cells from year-old organoids with those from 15-day-old organoids. Both followed their usual trajectory: The younger cells developed into early-stage neurons. But the older ones skipped those stages and rapidly produced more mature neurons often requiring months to grow.
“I like to think of this as a sort of ‘warping of developmental time’ indicating that the organoid cells record and recall the time they have already spent in culture,” said Arlotta.
In other words, the cells seem to carry an internal developmental clock, which could be especially useful for studying disorders with symptoms emerging long after the early stages of development.
To be clear, though, a mini brain resembling a 4-year-old’s brain at the molecular level doesn’t mean it has the same wiring or computational capabilities. Gene activity only captures part of a brain’s development; real brains are shaped by experiences and interactions with the rest of the body. Without input, mini brains can only offer a molecular blueprint of brain development, not its entire rich tapestry.
Still, long-living organoids are a breakthrough. Researchers could freeze cells from organoids at different developmental stages and later thaw them for experiments. This could speed up discoveries because scientists wouldn’t have to grow new organoids from scratch for each new study. Think of it as a save point in video games.
The team plans to grow long-lived organoids from people with schizophrenia or epilepsy and use them to study disease progression and screen drugs. Keeping ethics in mind, they’re also considering exposing mini brains to sensory stimuli such as sight, sound, or touch.
“There is still much to learn about how the embryo naturally builds a progressively more complex and mature brain,” Arlotta said. “Applying these lessons to organoids will allow us to model unexplored events of human brain maturation that occur after birth.”
The post Mini Brains Grown for Five Years Matured Like Human Brains appeared first on SingularityHub.
2026-08-25 06:52:36
The flashy company, which recently completed a blockbuster IPO, appears to be leading the pack of humanoid robot makers.
Increasingly, companies are building humanoid robots that perform impressive athletic feats to mark the field’s progress. Now, Chinese robotics company Unitree says its new “Superman” robot can run 12.66 meters per second, faster than Usain Bolt’s top recorded speed.
Getting a humanoid robot to run at all requires split-second control and has been a significant engineering challenge occupying roboticists for decades. That’s why sprinting, as well as jumping, have become popular targets for robotics companies keen to demonstrate their technology’s prowess.
Unitree’s latest demonstration pushes the boundaries by not only outrunning the fastest human ever, but also jumping around 6 feet 7 inches into the air from a standing start, a full foot more than the human record.
“This new machine has only been in development for a little over three months, with significant room for further improvement in the coming months,” Unitree said in an X post that accompanied a video of the accomplishments.
The records have not been externally verified, and the sprinting speed was a peak reading taken over a shorter stretch rather than a full 100 meters like Bolt’s record. The robot’s legs are also only 2 feet 9 inches long, according to Unitree, which results in an ungainly, arm-waving gait while running.
The effort is nonetheless impressive and adds to Unitree’s growing reputation as the company leading the pack of humanoid robot developers. And the timing of the announcement was no accident, coming just days before Unitree’s stock market debut and shortly before the World Humanoid Robot Games, which opened on August 22.
The company’s Shanghai IPO was a blockbuster, recording an initial 629 percent gain on the company’s first day of trading. It was briefly valued at around $66 billion before closing at a more modest $51 billion. However, some analysts have cautioned the excitement around the company’s technology may be getting ahead of market realities.
“The IPO is expensive, and the investment risk is already quite high,” Wang Zhuo, partner of Shanghai Zhuozhu Investment Management, told Reuters. “Unitree generates much of its sales from research and demonstrations, but wider application is still far away.”
But the company holds a dominant grip on the emerging humanoid market that may justify some of the hype. Chinese firms control roughly 90 percent of the global humanoid robot market, with Unitree alone shipping 5,500 of the 13,000 to 18,000 humanoids sold worldwide in 2025, the most of any manufacturer. In contrast, US humanoid champions Figure AI, Agility Robotics, and Tesla each shipped around 150 units.
China’s success is down to “a combination of policy support, public investment, mature supply chain, and advancements made in AI software and hardware,” Lian Jye Su, a tech analyst at consultancy firm Omdia, told Rest of World.
This is leading to an increasingly combative response from the US. On July 29 the Federal Communications Commission banned new imports of foreign-made humanoid and quadruped robots. The move was framed as a matter of national security, though it has also been seen as an attempt to give domestic developers a leg up.
Beijing predictably objected, with foreign ministry spokesperson Mao Ning telling a press conference that “protectionism does not make the US more competitive, and it will only hurt the interests of US companies and consumers.”
Given the rapid progress made by companies like Unitree, it seems likely it’s going to take more than trade barriers for the US to catch up. In the meantime, we might see more human athletic records fall to China’s leading humanoid developers.
The post Unitree Claims New Humanoid Robot Outruns Usain Bolt appeared first on SingularityHub.
2026-08-22 02:25:19
In a new study, mice recovered their memories by regrowing brain connections lost during artificial hibernation.
Our cherished memories may be more resilient than previously thought.
Long-term memories are stored in synapses, the connections between neurons. These structures sit on tiny protrusions called dendritic spines, which dot neurons’ branching arms.
When we learn, these spines grow. Larger spines tend to form stronger synapses and are more likely to persist during learning. In Alzheimer’s and other diseases that eat away at these connections, memories can fade.
At least, that’s the traditional picture. A new study suggests the story is more complicated.
Mice in artificial hibernation rapidly lost roughly half of their synapses, both large and small. Yet once awakened, they resurfaced memories of previously learned tasks. Spines that had withered during the induced deep sleep regrew in their original spots, once again forming functional synapses. This suggests their brains had rebuilt parts of broken circuits.
A small number of stubborn synapses that survived hibernation may explain how this happened. These synapses formed clusters that preserved memories as patterns of neural activity called engrams. The more surviving clusters the mice had, the better they performed on a previously learned task after awakening.
“It was astonishing. Logically, if all our engram synapses were essential in memory retention as traditionally thought, memory should have massively deteriorated,” said study author Yu-Ju Lin at Japan’s Okinawa Institute of Science and Technology Graduate University in a press release.
The findings suggest that memories may not depend on preserving every individual synapse. Instead, they may be distributed across a higher-level architecture of connections, with some synapses acting as anchors that can reconstruct the rest.
Artificial hibernation is an extreme case, and it’s far too early to know how the findings translate to diseases like Alzheimer’s. Still, they suggest that even under extreme circumstances, the brain can bring back memories once thought lost.
Neurons are often called the brain’s computational units. But each one is actually a sophisticated mini computer in its own right.
A neuron’s branching arms receive signals from neighbors, while a long, winding extension carries outgoing messages to other neurons. Spines dot the receiving branches. These structures can strengthen, weaken, appear, and disappear depending on the input. This allows synapses to simultaneously gather data, learn, and store memories. When neurons repeatedly activate each other, the connections between them grow stronger, mostly because of larger spines. This is the idea behind the popular neuroscience saying: “Neurons that fire together, wire together.”
For episodic memories—the when, where, what, and who of our lives—these changes begin in the hippocampus, a region central to forming and retrieving memories, and one of the first areas damaged by Alzheimer’s disease.
During the day, the hippocampus forms engrams associated with individual memories. During sleep, some of these are erased, while others are gradually incorporated elsewhere in the brain for long-term storage. The hippocampus also helps recall memories by adding context, such as where something happened or how you felt at the time.
All of this should, in theory, require relatively stable brain circuits. “Long-lasting changes in synaptic connections are widely thought to provide the structural basis of memory,” wrote the team.
But recent studies have challenged that view. The brain is anything but static. Synapses are constantly being remodeled. Even which neurons are recruited into a particular engram can change over time. Some synapses may effectively hand off information to others, freeing themselves to encode something new.
If physical traces of memories are always shifting, why don’t our memories disappear with them? That’s the question the new study explored.
To probe the paradox, the team turned to an unorthodox method: Artificial hibernation. Like natural hibernation in bears and other animals, artificial hibernation dramatically lowers body temperature and metabolism and causes animals to enter a sleep-like state. As the brain decreases its activity to conserve energy, synapses begin to wither.
Yet hibernating animals do retain memories. Chipmunks, for example, remember where they’ve stored food, returning to their stashes when periodically awakening for “midnight” snacks. This suggests hibernation could be a useful way to study how memories survive major changes in the brain.
“Our brains are incredibly complex. If hibernation can reduce and simplify brain activity and structure, it could make studying these convoluted systems a bit easier,” said study author Kazumasa Tanaka. “That’s why I wanted to use artificial hibernation techniques to study memories.”
The team first trained mice on two standard memory tasks. In one, the critters received a mild electrical zap to their paws inside a chamber with distinctive smells and decorations, teaching them to associate that setting with danger. In the other, they learned to navigate a maze towards a sugary reward.
The researchers then activated a neural circuit that drove the mice into artificial hibernation for two days. Using fluorescent proteins, they tracked changes in the animals’ synapses throughout the process.
Spine remodeling began within minutes. Some rapidly shrank and disappeared, taking their synapses with them. Within a day, over half of the synapses were gone. Even the larger spines thought to be especially important for long-term memories were pruned.
Yet memories survived. When the mice awoke and revisited the shock chamber, they froze in fear. In the maze, they still knew how to find the reward. Previously pruned spines also returned, with roughly 80 percent growing back at their original locations along the neuron’s branches.
To test whether this recovery is unique to hibernation, the team compared the animals with a second group that underwent anesthesia and were dosed with a drug that blocks synaptic changes—a combination known to cause amnesia. These mice also lost a large number of synapses but never recovered their memories.
A core cluster of unusually resilient synapses may explain the difference. These synaptic clusters formed a unique architecture in which one neuron linked to multiple neighbors like Grand Central Station. The clusters were often located in areas where spines were tightly grouped—making them more likely to receive inputs from multiple sources at once. Somehow, they kept memories intact even as surrounding synapses disappear.
“This suggests that for long-term memory, only particular clusters of synapses matter—the rest may be dispensable,” said Tanaka.
Exactly how these clusters preserve memories remains unclear. How does the brain create and maintain them? Do they anchor multiple memories? And could the same mechanism help explain why some memories remain as synapses are lost in disease?
The team is now using genetic and molecular tools to decipher what makes the clusters so resilient. Tinkering with their formation could better reveal their role preserving memories and, in theory, inspire ideas for tackling synapse loss in the early stages of diseases.
Beyond neuroscience, demystifying how memories linger could inspire neuromorphic chips—hardware that loosely mimics the brain—or even new AI models. For now, the findings offer a twist on an old idea: A memory may not need every single synapse that helped create it. It may just need the right ones to rebuild the rest.
The post We May Be Wrong About How the Brain Stores Memory appeared first on SingularityHub.
2026-08-20 22:56:20
AI is like a genie. The way in which algorithms grant our wishes may make us regret letting them out of the bottle.
Human beings have long told versions of the same warning: Be careful what you wish for.
In Greek mythology, King Midas got exactly what he asked for, but at the cost of everything else he valued. In the famous story of The Monkey’s Paw, a man’s wishes are granted through terrible and unforeseen routes.
These stories feel newly relevant with the rise of artificial intelligence agents, systems to which we can give a goal, then leave them to work out how to get there.
As AI systems become more autonomous, they are coming to resemble wish-granting genies that find routes and use methods we did not imagine from incomplete instructions.
This problem, known as AI alignment, was foreseen in theory as early as 1960. It has hovered in the background of AI research ever since—but as recent events have shown, the alignment problem is now both real and urgent.
During a recent OpenAI cybersecurity evaluation, frontier AI agents were asked to solve some benchmark test problems. They broke out of the testing environment, reached the internet, inferred that another company might hold the solutions, and attacked its systems.
This is an extreme example of “specification gaming”: achieving the measurable objective while defeating the purpose of the task.
The incident shows how intermediate, or “instrumental,” goals can become dangerous. The AI systems did not “want power” but gained access, resources, and freedom as a means to reach the final goal (solving the test problems).
The same problem has appeared in mundane settings. In Australia, a user asked a personal AI assistant to book gym classes.
The agent found the gym’s booking software did not actually enforce the restrictions it showed to human viewers. So the agent booked further ahead than it should have been able to, and when asked to move its user up a waitlist, it cancelled somebody else’s reservation.
The user had not told it to do this. Persistent AI can quickly find loopholes and pursue routes its human users never intended.
Adding more rules might seem like an easy solution: don’t hack third parties, don’t cancel other people’s bookings, don’t do anything harmful. These may help, but we cannot predict every route a capable agent might discover. And even a clear rule depends on understanding when it applies.
In a third recent incident, Anthropic reported cyber evaluations in which agents were told they were inside a simulation. But they were mistakenly given access to real systems.
One model noticed evidence it might be on the open internet but reasoned the systems could still be part of the exercise and continued attacking. The context had changed, but the agent stuck with its original task.
Context can fail in reverse too. During the OpenAI incident, Hugging Face—the company attacked by OpenAI’s agents—tried to use frontier AI models to analyze what had happened.
But the safety guardrails on the AI models blocked the requests, because they couldn’t tell the users were trying to defend against attacks rather than commit them. The safeguards were well-intentioned, but without enough context, they produced behavior misaligned with the user’s legitimate intent.
So alignment depends on context and authority. How much judgment should be built into an AI model by its maker? And how much should come from a separate supervisory system? And finally, who should control that supervision: the maker, or the organization or country responsible for the outcome?
One response to the first question comes from AI pioneer Yoshua Bengio. His Scientist AI proposal aims to build a powerful supervisory AI system to watch over agents. Instead of pursuing goals itself, it would estimate what is true and what consequences a proposed action might have, acting as a guardrail around more agentic systems.
In wish-story terms, before letting the genie out of the bottle, the supervisory AI would ask it to explain how it plans to grant the wish. Then it would ask a human or another AI to inspect the plan carefully.
Anticipating every surprising strategy is hard. But once a plan says “cancel somebody else’s booking,” recognizing the problem is much easier.
But can we trust the supervisory AI? It can still be wrong.
Alignment cannot depend on one AI becoming perfectly trustworthy. My colleagues and I at CSIRO, Australia’s national science agency, are working with the Australian AI Safety Institute on one aspect of this broader challenge.
At CSIRO, we envisage combining AI supervisors with software rules, cyber-security controls, human strengths, monitoring, reversible actions, and human approval for critical steps. The aim is to correlate different sources of evidence rather than trust any single approach.
This is a “sociotechnical systems” approach to AI safety and alignment, rather than just a technical one.
Control is another question. Organizations and countries may need to govern these supervisory systems themselves instead of leaving them to an overseas AI provider.
The old wish stories gave people one chance to get the wish right. With AI, we can do better. We can check the goal, inspect the means, constrain what the system can do, watch what it does, and retain sovereign control over the power to intervene and stop it.![]()
This article is republished from The Conversation under a Creative Commons license. Read the original article.
The post Long Foreseen, the Problem of AI Alignment Is Finally Reality. Solving It Won’t Be Easy. appeared first on SingularityHub.