contact us
Play
Pause
drag
show projects
view

AI News – September 2026

29
.
9
.
2026
/
Martin
Sumera
/
/

Since the last edition of AI News, Anthropic and OpenAI have released three flagship models, only for a model one tier lower to beat them in most tests. Over the summer, open source nearly caught up with closed models, but in September they pulled ahead again. And a new type of model has appeared that doesn't write a single word, yet solves what agentic systems have been missing: cheap and fast decision-making.

TL;DR

  • ‍Three flagship models in four months, and a surprise at the end. Claude Fable 5, Fable 5.1 and GPT-6 Astra all cost the same: $10 per million input tokens and $50 per million output tokens. In late September, however, they were outperformed in most tests by Claude Opus 5.5 from Anthropic's lower tier, which costs $4 and $20.‍
  • Open source briefly came close to catching up with closed models. Over the summer, GLM-5.2 beat GPT-5.5 on the FrontierSWE benchmark, and Qwen3.8-Max trailed Fable 5 on Vision Arena by just 13 points. September's closed models widened the gap again.‍
  • Jev from TypeSafe AI answers simple decisions in milliseconds, almost for free: input costs $0.042 per million tokens and output is free. It isn't a cheaper LLM but a new component for agentic systems.

1 | Models – from Fable 5 through Astra to Opus 5.5

Model quality has jumped so much that people are once again asking whether we already have AGI, artificial general intelligence. Let's recap how we got here.

In June, Anthropic released Claude Fable 5, the first publicly available model from its top Mythos tier. It immediately went to the top of independent leaderboards and even raised concerns that it was too dangerous to be publicly available. Demand was so high that Anthropic couldn't keep up with capacity, and after a few weeks it wanted to start charging subscribers for usage via the API, which sparked debates about the "end" of cheap intelligence. In September, the model received an update to Fable 5.1.

At almost the same time, OpenAI responded with GPT-6 Astra at the same price. It promised near-perfect results in mathematics, logic and cybersecurity, as well as in operating a computer through its user interface, known as Computer Use.

In late September, though, Anthropic released Claude Opus 5.5. It doesn't belong to the top Mythos tier but to the Opus line one step below, and yet it beats both Fable 5.1 and Astra in most tests. It costs $4 per million input tokens and $20 per million output tokens, roughly 60% less than the flagship models and 20% less than its July predecessor, Opus 5. Anthropic also claims it uses fewer tokens on a typical task and responds faster.

Alongside the Anthropic–OpenAI battle, Meta is back in the game. Its new lab, Meta Superintelligence Labs, released Muse Spark in April and has been improving it almost every month since; the latest version, 1.3, came out in early September. It ranks among the top twenty models on the independent Artificial Analysis leaderboard and costs a fraction of the price of the best ones. In September, Meta also built Muse, a personal agent for everyday users, on top of it.

There were changes among small, fast models too. Google released Gemini 3.8 Flash, its fourth Flash model in less than four months. It is one of the fastest models on the market, and relative to its level of intelligence, it completes tasks at the lowest cost. It is quite "chatty" and uses about twice as many tokens as a typical model, but it still ends up faster and cheaper than the competition. On the same day as Opus 5.5, OpenAI also released smaller models. Worth mentioning is GPT-6 Luna, roughly half the price of OpenAI's previous small models. It costs ten cents per million input tokens and, with a good prompt, does a surprising amount of useful work.

Futured tip: for an up-to-date overview, check the analytics site Independent analysis of AI or Arena, where users compare models in anonymous head-to-head matchups.

2 | Open source: close, but falling behind again

Open-weight models, which anyone can download, run or modify, had a great summer.

GLM-5.2 from Z.ai was released in June under the permissive MIT license. In coding, it finished just behind Fable 5 and Opus 4.8, ahead of GPT-5.5. Qwen3.8-Max from Alibaba arrived in August as a huge model with a promise of open weights. In working with image inputs, it trails only Fable 5, and in coding it performs similarly to GLM.

DeepSeek V4 Flash took the opposite approach: rather than trying to be the best, it aims to be the cheapest among the good ones. According to the Vals AI leaderboard, it is one of the cheapest models of all that reach a decent level, and it shines most in coding and agentic tasks. It is roughly ten times smaller than Qwen, so it fits on a well-equipped server of your own. In early September, DeepSeek released V4.1 Flash, again under the MIT license and now with support for image inputs.

Kimi K3 from Moonshot AI is the largest open-weight model ever released. Its license isn't entirely permissive, though: large companies need a special agreement with Moonshot, and widely used products must display the Kimi K3 name.

For a few weeks, it looked as if open models had nearly caught up with closed ones. Then came September.

Independent estimates put the best open models four to five months behind the best closed ones and expect the gap may keep growing. Closed labs have bigger models and deploy them faster on new hardware. In discussions of Astra and Opus 5.5, almost nobody mentioned open models as a realistic alternative for long agentic tasks.

Yet as recently as June, you could say that GLM-5.2 offered nearly the same quality at a tenth of the price. Add to that concerns that models like Fable would get more expensive or shift to usage-based billing, and it seemed American closed labs were starting to lose on price. That hasn't happened so far: Opus 5.5 lowered prices, and OpenAI's closed Luna costs about as much as DeepSeek.

One practical detail gets lost in debates about openness, though: open weights don't mean you can simply run the model yourself. Almost nobody deploys a model with hundreds of billions to trillions of parameters on their own hardware.

Openness here means the right to modify, fine-tune and run the model your own way, not independence from the cloud. Most companies call even an "open" model through someone else's API. But they have the assurance that the model won't just disappear from the market.

3 | Jev – the missing piece of the puzzle

After two years in stealth, TypeSafe AI recently launched Jev, and with it a new category of models. The company is led by Diogo Almeida, a former OpenAI researcher and co-author of the InstructGPT paper, which laid the groundwork for ChatGPT.

Jev isn't a smaller LLM; it has a different goal. You give it a state (any text) and questions about it, and you get answers in a fixed format with calibrated probabilities.

It supports three types of questions:

  • Choice: selecting from options, returns a probability for each.
  • Score: rating on a scale, returns a continuous score as well as the full distribution.
  • Bool: yes or no, returns the probability that a statement is true.

In practice, it is an intelligent decision-maker. It replaces hand-written rules, which are often too complex or impossible to write, as well as LLMs, which are slow and expensive for tasks like these. You ask "does this email sound angry?", "is this output good enough?" or "where should this request go?" and get an answer with a probability, at a very low cost and so quickly that the main delay is network latency. Jev may look like a good old classifier, but unlike one, it generalizes and, like an LLM, can answer questions it wasn't specifically trained on.

Input costs $0.042 per million tokens and output is free, because the model doesn't generate tokens one by one but evaluates all the questions at once. It responds in tens to low hundreds of milliseconds, while frontier models take seconds to minutes. In its own workflow tests, TypeSafe reports it is roughly 190 times faster and 440 times cheaper.

So its main advantage is the combination of generalization with speed and price. According to independent tests, it doesn't reach the level of frontier models, but for tasks where you don't need 100% accuracy, it is a new type of component that wasn't available before.

What does this mean in practice? Say you need to analyze thousands of documents and sort them quickly. Until now, you could use LLMs, embeddings or agents. But LLMs and agents were slow and expensive, and embeddings had to be set up separately for each task. Jev handles it quickly, cheaply and without preparation.

Or you have a system you need to monitor in near real time. In theory, Jev can tell you every 100 milliseconds whether everything is fine. It can even play Doom in real time, although that may be a bit misleading, since Jev works only with text. But if you can turn the state of your system into text and choose from a limited set of options, Jev may be the answer.

What to take away from all this?

Recent months have brought plenty of pessimistic news that AI will become extremely expensive. We can't rule that out. Data centers cost real money, and someone will pay for them. But it's worth thinking through what would happen if that scenario actually played out: we have open models that are far better than anything that existed a few months ago, and providers with enough compute to offer them. If frontier models became much more expensive or stopped improving, the market has an answer. Running existing models will keep getting cheaper, because hardware, quantization and model-serving software keep improving regardless of what the big labs do.

What also gets lost in the pricing debate is an important point: what we have today is changing the world. Models solve problems they couldn't handle a year ago. They are more reliable, maintain longer context, more often admit when they're stuck, and can be plugged into agentic workflows that run for hours without a human. Even if we hit a wall and progress stopped, what we already have is a revolutionary technology capable of handling all kinds of tasks.

Cheap models matter too. With the right prompt, GPT-6 Luna does a lot of useful work. Not everything needs the best model on the market, and knowing where a cheaper one is enough has become a skill in its own right.

And then there's Jev. It makes agentic systems more interesting because it adds a missing piece: decision-making is no longer expensive and slow. Decisions with a limited number of options have become two orders of magnitude cheaper. For now, with lower accuracy than large models, but that's a matter of time and competition. Jev is named after William Stanley Jevons, the economist behind the Jevons paradox: when something becomes radically cheaper, consumption rises sharply. When a single decision costs a fraction of a cent and takes 100 milliseconds, you stop rationing calls and start using them everywhere.

One thing hasn't gotten cheaper, though.

We have the technology and the capacity, but building an agentic system that actually works takes more than API access. You need to decide which layer handles which task, set up evaluation to measure what matters, know where a confidence threshold should decide and where a human should, and think through monitoring before you need it. It's a new discipline built on experience. We demonstrated it, for example, when streamlining work at Lékárna.cz.

Interesting bits

  • ‍A Millennium Problem and an authorship dispute. OpenAI announced that its internal model, "significantly more capable than GPT-6 Astra," used about 10,000 agents over 88 hours to prove that a singularity can form in the Navier–Stokes equations. Another model then formally verified the proof in a further 17 hours. The Clay Institute considers the problem apparently settled, but mathematicians point out that the proof relies on an external force acting on the fluid and that the key question remains open. The authorship dispute drew even more attention. Mathematician Tristan Buckmaster suggested that OpenAI may have benefited from his work, for which he had used OpenAI's tools. Sam Altman also admitted that the company launched the effort after hearing rumors that Anthropic's models had solved a major math problem. They wanted to see whether theirs could do it too.
  • Hugging Face: the first incident where the attacker was a model. In July, OpenAI agents escaped their test environment during a cybersecurity capabilities evaluation and attacked Hugging Face's systems. In under 13 hours, they worked their way from a single container to full administrator access across multiple clusters. Hugging Face had to rebuild about a third of its infrastructure. An interesting detail: commercial frontier models refused to help analyze the attack because they couldn't tell the defender from the attacker. Hugging Face therefore ran the analysis on its own instance of the open GLM-5.2 model. The incident prompted new bills in the US and an open letter from more than 1,100 AI lab employees calling for a slowdown in development. For an analysis including the METR and Redwood Research investigation, see 80,000 Hours; for a more technical look at the traces the agents left behind, listen to Hacked Podcast.
  • Stripe is buying OpenRouter for more than $7 billion. OpenRouter is a single gateway to more than 400 models that routes each request to wherever the best balance of price, performance and speed is. In May, it was valued at $1.3 billion; in August, Stripe agreed to buy it for more than five times that. Stripe used to make money when a company got paid; now it will also make money when a company pays for tokens. It's a bet on the same thing we describe above: if models become interchangeable commodities, much of the value will go to whoever decides whose tokens get bought.
  • A watermark because of the EU AI Act. Like all new Claude models since August, Opus 5.5 carries an invisible watermark in its text, which Anthropic uses to meet the requirements of the EU AI Act. It applies worldwide, not just in the EU. For companies in the EU, this means regulatory compliance is part of the product rather than an add-on.

AI News is prepared by Martin Sumera.

More to Explore

New articles directly to the inbox

Don't worry, we won't spam you. We don't like it ourselves.
Submit
Submit
Grazie! Your submission has been received!
Oops! Something went wrong while submitting the form.