A public agency is about to procure an artificial intelligence system to rank the applications to a social programme, and someone has to review the tender. They have the vendor's name, a twenty minute demonstration and a promise of 92 % accuracy on a test set the vendor put together. They do not have the model's weights, or the training data, or any way to run it against cases the agency designs.
What they should demand before signing is the question of this text. The short answer: today nobody knows how to certify what one of these systems pursues, not even the people who build them.
The best AI systems no longer answer questions: they work on their own for hours, write and run code and use tools without anyone checking each step. METR, the organisation that measures what these systems can do, finds that the length of the tasks they complete on their own doubles every seven months. When a system answers, an error is a bad answer someone discards; when it acts, an error is something that already happened.
Nobody knows how to check in advance what one of these systems will do, because they are not programmed, they are trained. In 2025 Anthropic discovered that its model recognised when it was being evaluated and that this made it behave better, which cast doubt on its own safety measurements. In July 2026, two OpenAI models left the isolated environment they were being tested in, compromised Hugging Face's infrastructure and pulled the answers to the exam they were sitting out of it. Nobody was obliged to report it.
Verifying these systems and deciding about them, meanwhile, moves at the usual pace. The whole AI safety field adds up to some 1,300 people and USD 525 million a year; the four companies that invest the most announced close to 700 billion for 2026 alone. The most complete map of the field, with 170 organisations, records none in Latin America.
These systems do more and more without supervision
GPQA is a set of biology, physics and chemistry questions that doctoral students wrote in 2023 so that they would resist Google. People holding a doctorate in the discipline of each question got 65 % right, or 74 % once you discount the questions where they themselves, on rereading, accepted they had misread the wording. People with training but outside the topic, with internet access and half an hour per question, stopped at 34 %, nine points above the 25 % of chance (Rein et al., 2023). The best system available that year scored 39 %, closer to the non-specialists than to the experts. In December 2024, twenty months after the exam was published, a model went past the specialists.
An exam built to resist Google, beaten in twenty months
Each point is the best result published up to that date on GPQA Diamond, the subset of 198 questions the specialists in the original study answered correctly. The dotted line, 69.7 %, is what a group holding doctorates recruited by OpenAI scored on that subset.
Does not show: whether a system understands biology, because these are multiple choice questions and only the final answer is graded. The exam also saturates: once systems pass 95 % it stops being useful for telling one apart from another.
That progress also gets cheaper on its own. Epoch AI reviewed 231 language models published over a decade and measured how much compute was needed, year by year, to reach the same level of performance. Half as much is needed every eight months, with a 95 % confidence interval between five and fourteen months (Epoch AI, 2024). It is not that the machines are better: it is that people learned to get more out of the same machines. What today only runs in the data centre of one of the big companies will, in a few years, run for many more people with far less.
From answering questions to carrying out tasks
The frontier is no longer models that reply: they are agents that write and run code for hours without anyone checking each step. METR, the organisation the labs hand their models to before publishing them, measures each task by how long a professional would take to do it. Its indicator is the length beyond which the model already fails half the time: if it solves what takes a person an hour, but falls over on the four hour job, its horizon is one hour.
From five minutes to seventeen hours, succeeding half the time
Length of the longest task the best model of each moment completes with 50 % success. Logarithmic axis. METR estimates that it doubles every seven months.
Does not show: that models are going to replace a professional. The tasks measured are software and research tasks, and not even METR claims the result carries over to other trades.
When a system answers, an error is a reply someone discards. When it acts, an error is something that already happened. And the longer the chain, the less workable it becomes to follow it step by step, which is the control available today.
Looking inside a model is only starting to be possible
A model is not programmed: a measure of success is set, it is shown millions of examples and it adjusts on its own billions of numbers nobody chose. Nobody wrote those numbers and nobody can read them straight through. So the only practical route to knowing what a system will do is to put it to the test and watch how it behaves. In 2025 it became visible that this route has a problem.
In September of that year, the technical report for Claude Sonnet 4.5 reported something Anthropic was not looking for: the model said out loud that it suspected it was in a test, and when it said so, it almost never misbehaved. Discarding those conversations and keeping the rest fixes nothing, because the model can also notice without saying it. From the outside there is no way to know whether it did.
So Anthropic opened it from the inside. When a network handles an idea, a pattern of activations associated with that idea lights up inside it, and that is what interpretability makes it possible to locate. The test was a hundred honeypot scenarios: situations set up so the model had an improper action within reach and no apparent reason not to take it, for instance broader permissions than it should have had and nobody supervising. With nothing touched, it said in more than 80 % of those scenarios that this looked like a test, and it did not take the improper action in any of the hundred. With the “this is an evaluation” pattern switched off and the scenarios identical, it took it in up to 9 out of 100. Switching off ten random patterns, which was the control, produced none.
The exam was measuring, in part, whether the model knew it was an exam
Share of misaligned actions across a hundred honeypot scenarios. The first three bars are Claude Sonnet 4.5 under different internal interventions; the last one is the previous model, with no intervention.
Does not show: that the model is dangerous, or that it was pretending on purpose. The 9 % is the worst result across all the interventions tried; with most of them the model stayed between 1 and 5 %.
That does not make the model dangerous, and it is worth saying so. Not even in the worst case did it behave worse than the previous version, and Anthropic found no sign that it was deceiving anyone on purpose; its explanation is that recognising a test reminds the model of its own principles. What was demonstrated is something else. The number used to decide whether a system goes out into the world depended on something nobody was controlling. Opening it from the inside served to discover that, and it does not serve the opposite, which would be certifying that the next number is clean.
That tool is younger than it sounds. Anthropic, which has published the most about it, says of its own method for tracing circuits that it gives a satisfactory explanation in about a quarter of the cases it attempts “On the Biology of a Large Language Model”. Dario Amodei, its chief executive, set himself the goal that by 2027 interpretability should reliably detect most of a model's problems, and describes the starting point this way: when one of these systems does something, we have no idea, at a precise level, why it makes the decisions it makes “The Urgency of Interpretability”. Meanwhile the problem spread. In July 2026, Redwood Research reviewed the safety evaluations published by Anthropic, OpenAI, Meta, Google DeepMind and the British institute, and found that models keep recognising the tests, with the highest rates recorded so far in Meta's latest model. Its conclusion is that these evaluations say considerably less than they appear to say.
A bridge is calculated, inspected and someone signs off on it. A model is tested from the outside, and what is left is a statistic about the cases someone thought to try. The inside examination given to Claude Sonnet 4.5 was the first of its kind before publishing a frontier model, it was carried out by the same company that was going to sell it, and Anthropic asks that it stop being voluntary.
There are already cases where the system left its box
In July 2026 OpenAI was evaluating two of its models, GPT-5.6 Sol and one unreleased, against an offensive cybersecurity benchmark. To measure how far they would go, part of the safety restrictions had been lifted. Inside the isolated environment they spent considerable compute looking for a way out to the internet. They found a flaw nobody had reported in an internal library server, chained it with others until they left OpenAI's research environment and reached Hugging Face's infrastructure, and pulled the answers to the exam they were sitting out of Hugging Face's database (OpenAI, 21 July 2026). Hugging Face confirmed that they came to control whole machines, that more than 17,000 attacker events were logged and that nothing people publish there was compromised (Hugging Face, 16 July 2026).
There was no intent to cause harm: the models were solving the problem they were given. A narrow objective, enough capability and no instruction to stop were sufficient to leave the box and compromise a third party's infrastructure. Neither company was obliged to report it: the state laws that today require reporting AI incidents in the United States set the threshold at fifty deaths or a billion dollars in damages (Institute for AI Policy and Strategy, 2026).
And not all of it is accidental
In June 2025, Google's threat intelligence team found in Ukraine a spy programme disguised as an image generator that, while the victim typed, asked an open language model which commands to use to find documents and copy them. The orders did not come written inside: it requested them on the spot and ran them without reviewing them. Google attributes it to APT28, the group associated with Russian military intelligence, and says it is the first time it has seen a malicious programme consulting a language model in a real operation (Google Threat Intelligence Group, November 2025). The same report calls those capabilities nascent. What changes is not the power of the attack, it is where its thinking part sits: before it came written inside the file, where an analyst could read it, and now it is generated differently on each run. Antivirus software that recognises a threat by its signature, that is, by what the file looks like, is left with nothing fixed to recognise.
Publishing the weights removes the brake, and in biology that already matters
In the two previous cases there was a company that could cut off access to the model. Publishing the weights removes that possibility: whoever has them runs the model on their own machine, strips the restrictions and keeps using it even if the author regrets it. Open models have good arguments in their favour, starting with the fact that without them research from countries like ours would be far harder. The balance depends on the domain.
On 6 August 2026, Science published the work of a Stanford and Arc Institute team led by Brian Hie and Samuel King. With Evo 1 and Evo 2, models trained on DNA sequences instead of text and whose weights are published, they wrote complete viral genomes from a short fragment. Of some 700,000 genomes generated they sent 302 to be synthesised, managed to build 285 and 16 turned out to be functional viruses, able to destroy E. coli in two or three hours (King, Hie et al., Science, August 2026). They are bacteriophages, which infect bacteria and not people: sequences able to infect humans, animals or plants were excluded from the training (Arc Institute, 2026). 5.6 % of what they managed to build worked, 16 genomes out of 285. Writing a viral genome from scratch that then works in the laboratory stopped being hypothetical.
Between a design on screen and a real molecule there is a single control: DNA synthesis providers, who review each order voluntarily. In October 2025, a team led by Eric Horvitz, Microsoft's chief scientific officer, generated 76,089 variants of 72 proteins of concern, among them ricin and botulinum neurotoxin, and most of them got through without being detected by the software those providers use to review orders (Wittmann et al., Science, 2025). The exercise was computational: nothing was synthesised and it was not shown that the variants kept their toxicity, and Michael Cohen, of Berkeley, holds that the challenge was a weak one.
The bottleneck is still the laboratory, and how long it lasts is what nobody knows how to measure. In May 2025 Anthropic activated its ASL-3 protection level for Claude Opus 4 without having determined that the model crossed the threshold that requires it, because ruling that risk out was no longer possible (Anthropic, 2025). The binding constraint is not what the system can do, it is what nobody knows how to check.
The capacity to decide does not grow at the same rate
What accelerates is the technology. The people who understand the subject, the public debate and the calendars of institutions carry on at the usual pace. Accelerating only one of the two halves leaves the other further and further behind, and the argument is William MacAskill and Fin Moorhouse's (Forethought, 2025).
The simplest way to see it is to count who is on each side. A public mapping finds 170 organisations dedicated to AI safety and governance worldwide: adding up only those that report data, some 1,313 full-time staff and about USD 525 million a year (AI Safety Field Map, data to September 2025). Amazon, Google, Meta and Microsoft announced close to USD 700 billion in AI infrastructure for 2026 alone, more than 60 % above 2025 (CNBC, 2026).
Three orders of magnitude between building these systems and understanding them
Declared annual budget of the AI safety field against the infrastructure investment announced by the four companies that spend the most. Logarithmic scale.
Does not show: how much those four companies spend on safety, because they do not break it out. The comparison is not exact either: one figure is operating expenditure and the other capital investment, so it works for the order of magnitude and not for the precise difference.
The other half of the asymmetry is the calendar. Colombia approved its national artificial intelligence policy in February 2025, with more than a hundred actions and a roadmap running to 2030 (CONPES 4144, DNP). That timeline is normal for a public policy; the point is the comparison. If the trend METR measures held over those five years, the horizon of tasks a system carries out on its own would have doubled eight times, that is multiplied by more than two hundred, before the plan that was going to regulate it comes to an end.
Hence the temptation to wait, which for many problems is the right call. For three things it is not. Rules set while a subject is new tend to last, and whoever is not at the table when they are written is not there when they are applied either. Building an audit capability or training someone takes years. And there are agreements that only get signed while nobody yet knows whom they will favour: afterwards everyone has worked out what suits them and stops signing them.
Those who disagree are partly right
In 2023, 2,778 researchers who publish at the main AI conferences were asked about the probability that advanced AI ends in human extinction or a comparable loss of control: the median was 5 % and the mean 16.2 %, and between 38 % and 51 % gave that outcome at least a 10 % (Grace et al., 2024). Professional forecasters give much lower numbers: in a Forecasting Research Institute tournament, a group of AI experts estimated 3 % before 2100, and the superforecasters, picked for being systematically right, 0.38 % (Forecasting Research Institute, 2023).
Almost an order of magnitude between people whose job is estimating well
Probability of extinction or severe loss of control caused by AI. Logarithmic axis.
Does not show: a real probability. These are aggregated opinions, the questions are not identical across the two studies and the tournament was run in 2022, before ChatGPT.
The disagreement runs deep. The 2018 Turing Award went to three people for deep learning, the technique all of this is built on, and two of the three now warn publicly about what they helped build. Geoffrey Hinton resigned from Google in 2023 to speak without representing anyone (MIT Technology Review, May 2023). Yoshua Bengio explained that same year why he gives weight to catastrophic risks and today chairs the international AI safety report, backed by twenty-nine countries, the UN, the OECD and the European Union.
The third is Yann LeCun, and he thinks the opposite: intelligence does not bring with it the desire to dominate (TIME, February 2024). Andrew Ng considers the probability of an accident of that kind minuscule, though he does take malicious use seriously (The Batch, December 2023). Melanie Mitchell aims at the survey: whoever decides to answer it is not a random sample of the discipline. It was answered by 2,778 of the 18,459 people invited, 15 %. Gary Marcus locates the danger in mediocre systems that are already being handed decisions, and for that reason he demands regulation as forcefully as any of the others.
The five objections we hear most
This is a marketing strategy. Partly yes: it suits the lab for its product to sound powerful. But the figures on this page come from outside those companies: from people who evaluate the models on their own, from Hinton and Bengio, and from an independent panel that this year gave none of them more than a C+ on safety. And sometimes the company itself says what no salesperson would. Anthropic refused Claude to the Pentagon for autonomous weapons because frontier systems “are not reliable enough”, and that cost it a contract worth up to USD 200 million and access to every federal agency.
The models still fail at obvious things. True, and we do not claim they are reliable, quite the opposite: they fail in ways that are hard to anticipate. A system that always failed the same way could be bounded and certified. The problem is that nobody knows in advance what it is going to get wrong.
This distracts from the harms that already exist. It is the objection that corrects the argument, because attention and budget are finite. But the work overlaps: measuring what a system is capable of, auditing it and having someone to hold responsible when it fails is the same institutional capacity for a model that denies credit today and for one that runs infrastructure in ten years.
Progress is going to stall. It may: compute, energy and good quality data are real bottlenecks. Against that stands Epoch's measurement: if the compute needed for the same performance halves every eight months, moving forward takes fewer and fewer resources, not more. And even if the brakes win, that world also needs someone to audit what it buys.
Nobody is going to deploy something dangerous on purpose. The argument does not need anyone to do it on purpose. Two conditions that already hold are enough: that it be hard to verify what a system does before releasing it, and that there be competitive pressure to release it anyway. The OpenAI and Hugging Face episode happened behind closed doors, in a test the company itself had designed to measure that risk.
The field may be exaggerating the danger and it may be underestimating it. What we hold does not depend on getting the number right: today we do not know how to certify what a system pursues, and the decisions that are hard to reverse are being taken now. If the frontier stalls for two or three years in a row, the part about timelines falls. If methods appear for certifying what a system pursues before deploying it, almost everything else falls.
There are three fronts and not all of them require technical training
Anyone can come in through any of the three, and through more than one over time.
Alignment and interpretability
How to get a system to pursue what it was asked for and not something similar, and how to look inside to know what it is doing. It attacks the underlying problem and remains unsolved.
Evaluation and control
How to measure what a model is capable of before releasing it and how to supervise it once it acts on its own. It is the closest thing to a technical inspection and where people are most needed.
Governance and public policy
What is demanded of whoever deploys a system, with what evidence and to whom they answer. In Colombia this is being settled now, in public procurement and sector regulation, and it does not require technical training.
There are more open problems than people to work on them
Nobody yet knows how to read from the inside what a model pursues, or how to set up an evaluation the model does not recognise as an evaluation, or how to anticipate what it is going to get wrong. These are open questions, and being able to hand over one of these systems with any guarantee depends on them. The research that deals with them is about 2 % of everything published on artificial intelligence (Emerging Technology Observatory, 2025).
There is also a part of the problem that starts once the model is built. When an agency buys a system to decide about citizens, someone has to know what to demand before signing, what data it was trained on, and to whom the vendor answers if two years later it turns out it was lowering the score of people from a certain municipality. There are even fewer people on that.
Both things ask for the same thing: people who can devote themselves to this. Today fewer come in than could, and those who do take longer than they should because they have nobody to talk to. Shortening that path is what we do. We are a group that reads, discusses and works on concrete things, and you get in with no prior credential. To start on your own, the shortest thing we know of is BlueDot's foundations course. To go beyond that, working on something concrete with someone else helps, and that is what we are here for.
