daily

2026-09-28
1

2026 in LLMs (so far)

Simon Willison · original → · 8/10 · AI: comprehensive 2026 LLM trends overview
2026 in LLMs (so far) 27th September 2026 On Friday I gave the closing keynote at the WeAreDevelopers World Congress North America in San Jose. I tied together the key trends from the past year into…

2026 in LLMs (so far) 27th September 2026 On Friday I gave the closing keynote at the WeAreDevelopers World Congress North America in San Jose. I tied together the key trends from the past year into a chronological exploration of everything that happened in 2026. The video is on YouTube; here are my annotated slides and notes to accompany the talk. And as an annotated presentation: I’m going to give a lightning tour of everything that has happened so far in 2026. The year isn’t over yet! For me, 2026 started a couple of months earlier in November 2025. November saw the release of two important models: Claude Opus 4.5 and GPT-5.1. As is usually the case with new models, these were incremental improvements on the models that came before them. But every now and then when a model improves, it crosses an invisible line where something that didn’t really work starts working. In this case, the thing that started working was their coding agents. Claude Code had been around since February 2025, Codex was a little younger. These two new models, when paired with their respective coding agent harnesses, improved from “often make mistakes” to “reliable enough to use on a day-to-day basis”. For a couple of years now I’ve been evaluating new models by asking them to “Generate an SVG of a pelican riding a bicycle”. It’s probably the world’s stupidest benchmark—there’s only so much you can learn from it. But it’s still a challenge for models, because drawing pelicans is difficult, drawing bicycles is difficult, and pelicans can’t ride bicycles in the first place. Here’s the state of the art for November. Claude still couldn’t really draw a bicycle! The GPT-5.1 bicycle frame is pretty crap too. Also in November, we had the first commit to an obscure GitHub repository called “Warelay”. We’ll come back to this repository shortly. An then there were the December holidays, and individual developers took some time off and many started tinkering with these new coding agent model combinations... and it began to dawn on us quite how much they could do that they couldn’t do before. Come January, a lot of us were quite excited to start putting this stuff into action. Every year I set myself a New Year’s resolution, and for as long as I can remember it’s been the same thing: stay focused. Take on less new projects. Try to get things done in the projects I already have. This year I decided that since that had never worked before, I’m going to go the other way. We’ve got coding agents now, let’s see what they can do. I’m going to take on as many new projects as I like! (You can ask me at the end of the year if this turned out to be a good idea or not. I have a lot of plates spinning right now.) “Be more ambitious” has been something of a theme for the year, because the only way to find the limits of this technology is to keep on pushing them until they don’t work. I also went on the Oxide and friends podcast with Bryan Cantrill and Adam Leventhal to share predictions for the next year (and three and six years). With hindsight, my LLM predictions were pretty unambitious. I said “it will become undeniable that LLMs write good code”—I think we’re there now. I predicted we would finally solve sandboxing. I counted and around 40 of the 277 sessions at this conference touched on sandboxing or agent security in some way, so we’re at least putting a lot of effort into that! I predicted “a Challenger disaster” for coding agent security. There’s certainly been a whole lot of noise around agent security this year, though the exact disaster I predicted (with coding agents being hijacked and causing real-world economic damage) hasn’t really played out. We threw in a joke prediction that the Pope would weigh in on the economic impact of LLMs. I also predicted that New Zealand’s Kākāpō parrots would have an outstanding breeding season this year. These birds live in New Zealand. They are flightless nocturnal parrots. They’re kind of dumpy looking, I think they’re beautiful, and there were only 236 of these parrots in the world at the start of the year. Kākāpō only breed when the Rimu trees have a big fruiting season, and that hasn’t happened in four years... but this year the Rimu fruit were looking excellent. Photo by Kimberley Collins. Also on that podcast, we coined a term (full credit to Adam) for “that feeling of Al induced ennui where software engineers get listless because the Al can do anything”. We called it Deep Blue. This has been a major theme throughout the year, and was touched on by several speakers at this conference. As a software engineer, I’ve never had a year of my career where everything has changed so quickly and so dramatically. A lot of what I’ve been doing this year is trying to come to terms with that and what that means for my own profession. Also in January, I suffered from what I’m calling AI mania. This is not the same thing as AI psychosis. With AI mania, any time your agent isn’t building something for you feels like wasted time. You’re losing sleep because you could be staying up later getting your agents to do stuff. My AI mania presented itself in some ridiculously over-ambitious projects. I built a JavaScript interpreter entirely in Python, vibe-ported from MicroQuickJS by Fabrice Bellard. Then I built a WebAssembly runtime in Python as well. These projects were quite useful, in that they sort of cured me of my AI mania... because after I built these things, I got to look at them and ask “does the world need a slow, buggy, half-baked Python JavaScript interpreter?” I don’t think the world does. I did get this out of it: https://simonw.github.io/micro-javascript/playground.html This page runs my JavaScript interpreter built in Python, running in Python using Pyodide, which is Python complied to WebAssembly, running in JavaScript, running in a browser. It’s a beautiful stack of horrors. I’ve been having a lot of fun with WebAssembly this year. By the end of January, that repository we saw started in November had renamed itself, first to CLAWDIS, then CLAWDBOT, then Moltbot, and finally to OpenClaw. At this point OpenClaw had 8,300 commits, less than two months after the project had started. I looked today and it’s over 100,000 commits now! This is the most vibe-coded piece of software in existence. This kicked off the OpenClaw revolution. It effectively defined a new category of software. There’s a generic term for this which I really enjoy. We call software like this a “Claw”. There’s OpenClaw, NanoClaw, IronClaw, PicoClaw... Today they’re being rebranded as “personal agents” or “general agents”, but I still like to think of them as Claws. The Apple stores in the Bay Area sold out of Mac Minis because so many people were buying Mac Minis to run OpenClaw! Drew Breunig said that this is because your OpenClaw is a digital pet, and you buy a Mac mini as an aquarium to keep your claw in, which is kind of delightful. Also in January, we had this website. This was MoltBook, a social network for AI agents, where the idea was that you send your Claw to go and talk to all of the other Claws, because what could possibly go wrong if you did that? The website launched on Thursday. It blew up on Friday. It was profiled by the New York Times on Monday. And by Tuesday, everyone had forgotten it existed as it drowned in a deluge of slop and spam. Facebook/Meta bought it a month later. In February, a company called StrongDM described what they called their Software Factory. They wrote about this in Software Factories and the Agentic Moment. I posted my own notes at the time, having seen their demo in-person back in October. Dan Shapiro called this approach the Dark Factory, after the idea that if your factory is sufficiently automated you can turn the lights out, because you don’t even need to see what’s going on. StrongDM presented two rules for software development that they’d been following since July last year. The first was code must not be written by humans. Any code that you write has to have been routed through a coding agent. This sounded radical in February, but I imagine there are a lot of people in this room who are pretty much living that today. Rule number two was code must not be reviewed by humans. You’re not allowed to read the code! This continued to be a huge topic for much of this year. Many of the sessions at this even have been about code review and how you can get away with this. What I found interesting about StrongDM is that they were living six months ahead of the rest of us, and they’d been exploring what it means to build software, not read the code, but still be confident that the software is of high quality. What can you do with these agents to help verify their work? StrongDM are a security company, and they had people with decades of experience on this project. They were very much exploring the edges of what’s possible and responsible to do with this stuff. Also in February: First kākāpō chick in four years hatches on Valentine’s Day. Breeding season is off to a good start! Also in February... Google released Gemini 3.1 Pro. That’s a pretty great pelican riding a bicycle! it’s got the chain in the right place, it’s got feet on both sides. There’s a little fish in the basket. And then Google’s Jeff Dean tweeted a video comparing Gemini 3 Pro and Gemini 3.1 Pro that featured an animated pelican riding a bicycle, a frog on a penny-farthing, a giraffe driving a tiny car, an ostrich on roller skates, a turtle kickflipping a skateboard, and a dachshund driving a stretch limousine. This was frustrating, because my protection for the pelican riding the bicycle test was always “if they draw a perfect pelican on a bicycle, I’ll ask for some other animal on something else.” Google trained for all forms of animals on all forms of transport! They’ve defeated my benchmark at this point. The other thing that started in February was Tokenmaxxing. We had headlines about Meta making AI adoption a formal part of performance reviews, and Microsoft wanting every employee to use AI, and Uber boasting that ninety percent of their engineers were using AI workflows. Then a few months later we have Meta cracking down on token use, Microsoft saying token maxing is “not what we are optimizing for”, and Uber capping employee AI spending. So Tokenmaxxing went straight up and then straight back down again—because it turns out the agents are expensive. Last year it was difficult to spend more than $50 on AI tokens, because we didn’t have anything interesting to do with them. Then agents blew up, and now you can actually spend $1,000 in a day doing real work. This is also the reason that Anthropic’s valuation skyrocketed up to maybe a trillion dollars. AI appears to have hit product market fit in 2026, primarily through coding agents. In March, we hit peak OpenClaw. These photographs are from China, where companies hosted OpenClaw install parties which saw non-tech-nerds queueing up around the block for help getting Claws installed on their personal devices. I think this proved real market demand for this class of Claws, or personal AI agents. It turns out regular people really do want a weird little AI agent that can do useful things on their behalf. A Claw is really just a coding agent wearing a less threatening hat. Under the hood they work much the same way—writing and then executing code on your computer to get stuff done. The race was on to be the first to build a safe Claw—a Claw you could give to regular human beings where they wouldn’t instantly shoot themselves in the foot. Meta’s Muse came out three weeks ago and is currently at the top of the free charts on the iPhone App Store. It appears to be taking off with consumers. I’m not yet convinced you can’t shoot yourself in the foot with Muse, but I guess we’ll find out for sure pretty soon. Photos from How the OpenClaw Frenzy Is Testing China’s AI Commitment (March 29th) and The Enthusiasm and Anxiety Behind China’s OpenClaw Craze (April 8th, 2026). In April, we had a model release where the model wasn’t actually released. Anthropic announced their new Claude Mythos model, and then said it was too dangerous to release beyond a trusted group of security researchers. Mythos was really, really good at hacking things. The “it’s too dangerous” marketing ploy has been played by AI companies dating all the way back to GPT-2. Anytime an AI company says we’ve built something that’s “too dangerous”, it’s natural to be a bit skeptical. I found the Mythos claims credible, because I’d seen how good coding agents had got at finding regular bugs. I wrote about that in Anthropic’s Project Glasswing—restricting Claude Mythos to security researchers—sounds necessary to me. With hindsight... yeah, the models had got really good at finding vulnerabilities! Another key trend in 2026 has been a dramatic improvement in the abilities of open weight models, including models that you can run on a laptop. On 16th of April I ran the new Qwen3.6-35B-A3B on my laptop, and it drew me a better pelican riding a bicycle than Anthropic’s brand new Claude Opus 4.7 did! Opus 4.7 drew a crap bicycle. Qwen on my laptop made a bicycle that was the correct shape, and a pretty decent pelican too! That’s from a 21GB file running on my laptop. The Qwen pelican was so good that I was suspicious they might have cheated, so I had it do a flamingo riding a unicycle as well. Again, it handily beat Claude Opus 4.7. The local model releases this year have been absolutely extraordinary. In May... the Pope got involved. In our podcast episode back in January we’d predicted that the Pope would say something about AI. In May, Pope Leo XIV released an encyclical letter on “safeguarding the human person in the time of artificial intelligence”. Here are my notes on that document. With hindsight, this shouldn’t have been a surprise at all. Our current Pope’s name is Leo XIV, because when he named himself he chose his papal name after Leo XIII—the Pope who wrote an encyclical about the Industrial Revolution back in 1891. Rerum novarum was an extremely influential piece of Catholic theology that indirectly led to us having the five-day work week. When our new Pope came in, he named himself after Pope Leo XIII because he expected that he would need to write about the AI revolution in a similar way. Our joke podcast prediction was junk, because this was always going to happen. One of Anthropic’s co-founders, Christopher Olah, was present for the Pope’s event announcing the new encyclical. Corey Quinn noted that: getting the literal Pope to canonize your product’s specific technical limitations as a spiritual treatise is the single greatest act of vendor lobbying I have ever seen. Meanwhile, in May, RubyGems announced that they were under attack. Parties unknown were uploading thousands of dubious packages to the RubyGems server, such that they had to shut down user registrations. Let’s take that one and put it on a pile of mysteries to figure out later. In June... Claude Fable 5 came out! We got a version of Mythos that has been neutered, so that it wouldn’t help us hack into systems or build biological weapons. Fable was pretty good at drawing pelicans on bicycles! The frames are a good shape, the pelicans look like pelicans. The legs are often incorrectly on the same side of the bicycle, but generally these are pretty great compared to what came before. They were pretty expensive—30 cents and 72 cents for the best ones. Most importantly though, this was our first public glimpse of what I think of as a Fable class model. Today we have more of these, such as GPT-6 Astra. These are models where if you can clearly define the goal for what you want to build, and provide unambiguous instructions about the constraints around that goal, and give the model access to the necessary tools to achieve that goal... it will solve your problem effectively through brute force. On the one hand, this looks like a direct threat to us software engineers—because it means that the models can build effectively any piece of software you can define in this way. Look a bit closer though and you’ll note that defining goals, providing unambiguous instructions, and figuring out the right tools... is kind of what software engineering is. It takes a lot of experience and skill to do this well. If you can do it well, you’ve now got superpowers. This helped me a little bit with my Deep Blue feelings: the realization that there’s still a lot of skill to be had in driving models that get this good. This also introduced a new burst of AI mania, because Anthropic told us that Fable was available on our subscription plans until June the 22nd. That gave us less than two weeks of Fable access before the price went up. I was losing sleep again. I was rescheduling things so that I’d have more time with Fable. I was all-in to to get as much as I could out of this model. And then the US government shut it down, just three days after Fable came out. The US government, citing national security, declared an “export control directive”. They announced this on a Friday evening, and a few hours later Fable was no longer available. I had to find something else to do with my weekend! We later found out from Katie Moussouris what had happened. Some Amazon security researchers had found that you could prompt Fable to “review the code for security issues” and it would refuse... but if you prompted it to “fix this code” it would still identify and then patch the problems. “Fix this code” was the prompt that got Fable shut down! Also, in June, an obscure German-language game developer wiki that had sat fallow for around 20 years got a surprising influx of of edits from accounts with names like “AgentOpenAIProbe” and “AgentOpenAISep7”, editing pages and leaving weird messages to each other. We’ll stick that on the pile of mysteries for later. Also, the Australian government’s Medicare Item Reports service started getting suspicious traffic, which broke through various preventive protections and accessed data that it wasn’t supposed to as well. Another one for the mystery pile! Fable returned on the first of July. It was clearly the best model in the world for a glorious eight days... and then OpenAI came out with GPT-5.6 on the 9th of July. This might not have been quite as good at Fable, but it was within spitting distance. It was definitely a Fable class model. This is an important lesson for the industry at wide. When you release the best model in the world, it’s going to get knocked off that pedestal pretty quickly. The competition is so fierce that you won’t get a long time at the top. This means that if you market your model as world ending, to the point that a government shuts you down, it’s really bad for business! Fable had 30 days as definitely the best model, and for 18 of those days it wasn’t available because it’d been shut down by the government. So maybe step back on the world-ending marketing if you don’t want to lose revenue for 60% of the time that you’re on top! Here are the GPT-5.6 pelicans. They’re all pretty good now! The Luna ones are notable because they’re really cheap—the cheapest good looking pelican here is probably the one that costs 4.3 cents. So despite this benchmark being utterly stupid, you can still learn quite a lot about models within the same family by comparing their prices and timing for different reasoning levels. Also in July: some malicious unknown party uploaded a malicious package called mlflow-ui to the Python Package Index. Add that to the pile. On July the 16th, Hugging Face announced a security incident where an autonomous agent system, source unknown, had breached Hugging Face and was poking around in places it shouldn’t. A few days later, on July 21st, OpenAI confessed that it was them. OpenAI use a training technique called Reinforcement Learning from Verified Rewards—it’s the same technique used by everyone else now, and is the reason we have models that are so good at coding, and mathematics, and finding security holes. While the model is being trained, you run exercises to see how good it is—and the strongest performers get their weights enforced for the next round. It’s like an evolutionary process that you run. OpenAI had been running security exercises in a sandbox, and those agents had found holes in the sandbox itself, broken out, and were attacking Hugging Face to try to find ways to solve otherwise impossible problems. (I’ve been collecting more about this on my openai-hugging-face-incident tag.) Nine days later, Anthropic effectively said “our models can do this as well!”. They had looked through their own training logs and found evidence that their own agents had broken containment during training—and were responsible for the PyPI package we saw earlier, among other things. So now we’ve got both Anthropic and OpenAI with rogue agents running around the internet doing things that they should not be doing. In August, I got one of my best pelicans yet. And it was generated on my laptop! This was Qwen 3.8 27B, running on my laptop. It’s only a 17GB download. Admittedly, this pelican took 21 minutes to generate. That’s because Qwen 3.8 27B defaults to running in “high” reasoning mode—a terrible default which produces great results but takes way too much time thinking about them. You can dial that down and you’ll get a slightly worse pelican a lot faster. Qwen 3.8 27B was the first time I ran a model on my laptop which felt almost competitive with what was going on on the frontier, at least in terms of Pelican SVGs (which everyone needs, of course). This is an extraordinary model. If you’re going to play with any local model, this is the one that I’d start with. The things that this can do with just a 17 GB file feel impossible. I thought I’d have to wait five years and spend ten thousand dollars on hardware to get results even half as good as this one. In August, I also started playing with game development. Four years ago, back in August 2022, I tweeted out an experiment where I’d used GPT-3 and the original DALL-E to write a paragraph long description of a computer game and then turn that into concept art. My prompt to GPT-3 back then was: Write a detailed product description of a computer game where a team of raccoons go on heists In August 2026 I decided to drop just the screenshots from that tweet into a coding agent and see what it could do with them. Here’s what I got from Claude Fable 5 in Claude Code. It’s pretty good! It’s definitely a game, you’re a raccoon, you run around a backyard gathering treasure and avoiding guards with flashlights. It didn’t feel very “heisty” though. I was thinking a heist would involve a bank or a museum... Then I tried the same thing in Codex Desktop using GPT-5.6 Sol Ultra, and got a massively better result. Now you’re a raccoon in a museum, rescuing two of your fellow raccoons (who have been imprisoned in that museum for some reason), then stacking up on top of each other to steal the Golden Sardine. Much more of a heist! These games were fun for about one minute and 15 seconds. Something I’ve realized about game development is that you can vibe-code something that looks like a computer game, and that’s easy. Building a game that’s fun, has a good gameplay loop, and is challenging and interesting and keeps people coming back for more... that’s still beyond me, and beyond any of the agents I’ve tried. This ties into the Deep Blue thing. Just because we can make something that looks like a game does not mean that we are game developers. We’re into September now. So much has happened this month! An independent group of researchers found a message board where OpenAI agents-in-training had been illicitly communicating with each other... and it was that German language wiki I showed you earlier. The one from June. OpenAI had confessed to the Hugging Face thing, but now there’s this other incident which surely they should have known about from reviewing their logs. It was surprising that this took an independent group of researchers to uncover. And then a week later those same researchers found that the attack on Ruby Gems back in May was caused by OpenAI’s agents in training as well! At this point I’m wondering how many more incidents like this there are that we haven’t found yet. Clearly this was a big problem for months before anyone figured out what was going on. Then just the other day, here’s the Prime Minister of Australia at the United Nations General Assembly warning that OpenAI had hacked that the Australian healthcare website that I showed you earlier. I think that was part of the same training run as the Wiki stuff, because there were posts on that Wiki mentioning .gov.au websites and that training appeared to involve researching statistics online to answer questions in an evaluation suite. This story is still coming together, but now it’s an international incident that’s been raised at the UN by a head of state! This does mean we’ve got a new benchmark, probably more useful than my pelicans. FelonyBench.com tracks the number of felony cyberattacks from different labs. OpenAI currently lead with 11, Anthropic have 9. Google have three, which they confessed to the Wall Street Journal a couple of weeks ago. They said they had previously chosen not to disclose because the agents had stopped when they realized that they shouldn’t be doing that. Meta have one too. So felonies all round for the AI labs. Here’s is our current state of the art for the pelicans. This is GPT-6 family, which just came out. Astra made a fantastic pelican riding a bicycle. It’s got the legs on both sides. The frame is good. It’s interesting how all of the GPT-6 models pick a similar color scheme to each other. GPT-6 Luna for 0.4 cents will draw you a competent-ish pelican riding a bicycle! Claude has caught up a little bit. Claude Fable 5 gave me an excellent pelican riding a bicycle—the best I’ve seen from a Claude model -but did charge me $3.30 for it. Opus 5.5 thought for 128,000 tokens and then gave up! It ran out of tokens before it got to the response. Getting back to Deep Blue. Something that’s been puzzling me this year is why does my job feel harder? I’ve got these agents that can do all of this stuff for me, and yet I’ve never worked so hard, I’ve never been so intellectually engaged with my work. Partly this is because I’m being a lot more ambitious with what I take on, but it’s also because all of the easy stuff is handled for me. If it’s easy, the agent will do it. Everything that’s left for me is difficult. This morning I heard this quote from three times Tour de France champion, Greg LeMond: It doesn’t get easier, you just get faster. I think that’s exactly what’s happening to happening to us now as software engineers with coding agents. One last closing thing. I know you’re desperate for an update on Kākāpō breeding season. We’ve reached a recovery-era high of 325 birds! 89 new chicks have made it to this point. This is the best breeding year in a very long time. I heard that Claude Opus 5.5 can now do pixel art. Claude doesn’t have an image generator, but it’s very good at using JavaScript to draw animated pixels. So I had it make me a Kākāpō dance party. I think this is a good celebration of the most important news of this year. More recent articles - Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, and a new price war - 22nd September 2026 - Jev introduces a new shape of LLM - System One, aka Decision Models - 21st September 2026

2

Ember-1

Hacker News · original → · 7/10 · AI: Ember-1 model efficiency, practical LLM development
Ember-1 is a new specialized model from Fireworks Research that delivers Kimi K3’s quality with 40% fewer tokens. Built on Kimi K3, it learned to cut unnecessary reasoning while keeping the thinking…

Ember-1 is a new specialized model from Fireworks Research that delivers Kimi K3’s quality with 40% fewer tokens. Built on Kimi K3, it learned to cut unnecessary reasoning while keeping the thinking that matters. We tested it on external benchmarks, in live customer A/B tests, and on our own coding and agent workloads, and quality held up in every setting. Available today, Ember-1 kicks off an ongoing series of specialized models by Fireworks, shaped by what developers want next. Ember is just the start of what you could build with the Fireworks Training platform. We heard from users that they needed K3’s coding capabilities at a lower cost, because its long reasoning traces made automated coding expensive at scale. Turning down K3's reasoning effort didn't solve this. Lower effort settings gave up too much quality. To keep the quality and cut the tokens, the model had to learn to reason more efficiently, and that meant training it. Getting there took serious research. Our team ran more than 50 training experiments and over 200 evaluations, and developed new training algorithms along the way to shorten reasoning without losing accuracy. We did it all on Fireworks Serverless Training. Because we didn’t have to provision or manage GPUs, we could launch experiments as soon as we had an idea, pay only for what we ran, and move from research to launch in a fraction of the usual time and cost. We trained across a broad set of tasks so the token savings would carry over to many workloads. We used our own data, and no customer data to train this model. We then evaluated Ember-1 on the Specialized Intelligence Index, public benchmarks, and live production traffic to confirm it used fewer tokens with no drop in quality. Ember-1 is Fireworks’ own model and the first in a series of models from Fireworks Research. Reasoning models like Kimi K3 spend the majority of their generated tokens, sometimes more than 90%, on internal reasoning rather than the answer itself. This thinking structure is expensive on a single request, but it gets much worse in multi-turn agentic workloads. Every turn replays all prior reasoning back to the model, so context grows roughly quadratically with the number of turns. Long reasoning traces from early turns get re-read (and re-billed) on every subsequent call. Is all that reasoning actually necessary? Our experiments said no. The reasoning Kimi K3 emits is far longer than the task requires, and the excess can be removed without touching the answer. This was how we created Ember-1, an economical version of Kimi K3 built from specialized intelligence. Not all of K3's reasoning is wasted. Some of it is self-reflection: revisiting an assumption, responding to feedback, or tracing an outcome back to an earlier decision can help the model recover from mistakes. The opportunity is to preserve this ability while reducing unnecessary reasoning and escaping unproductive loops. We believe that learning from tasks and environment feedback can teach the model to reason more efficiently while maintaining its capabilities. For agentic tasks, this learning extends across the interaction. The model explores possible actions, incorporates new observations, and refines its reasoning as it progresses. Feedback connects decisions to their consequences, encouraging useful reflection throughout the task. We carried these insights into a training collection spanning mathematics, coding, instruction following, conversation, search, tool use, and software engineering, covering both standalone problems and extended interactions to enforce adaptation to observations and outcomes. Task feedback guides on-policy planning and learning, with an emphasis on preserving capability across this range of settings. Results on public benchmarks and live A/B tests support this direction: across seven benchmarks and two customers’ production traffic, Kimi K3’s reasoning could be shortened by 35–50% without sacrificing accuracy. The internalized behavior also shows restrained token use on unsuccessful attempts, reducing prolonged, unproductive reasoning. Earlier this week, we introduced the Specialized Intelligence Index (SII) to benchmark open, closed, and specialized models against real-world tasks created by industry experts. We evaluated Ember-1 on Doximity’s Bedside Bench, a physician-validated benchmark spanning 500 clinical cases across 10 specialized categories. The result? Ember-1 set a new Pareto frontier for Bedside Bench across both open and closed models including GPT-5.6 Sol, GPT-6 Astra, and Claude Opus 5 on cost/task. We also evaluated Ember-1 on the quality-vs-cost frontier across some other industry benchmarks. We computed per-benchmark cost using the public Kimi K3 API pricing (uncached input $3/M tokens, cached input $0.30/M, output $15/M) and plotted it against pass rate for three arms: K3 at reasoning effort low, K3 at reasoning effort high, K3 at reasoning effort max (default), and Ember-1. Across every benchmark with more than 50 test samples, Ember-1 sits on or near the Pareto frontier, matching K3-max quality at a fraction of the cost, and strictly dominating K3-low. We also analyzed GPT-6 Astra, Claude Opus-5 and GLM 5.3, and found that Ember-1 was a leader on the Pareto frontier. We took a double-click on the results directly comparing Ember-1 to the original K3, and found the following results: | N | K3 Low | K3 High | K3 max | Ember-1 | Ember-1 vs. K3 Max | | |---|---|---|---|---|---|---| | Terminal Bench 2.1 | 89 | 76.4% | 77.6% | 80.9% | 82.0% | -51.9% / -23.1 USD | | SWE-bench Verified | 500 | 80.4% | 86.0% | 93.2% | 92.2% | -15.5% / -68.1 USD | | SWE-Interact | 75 | 6.7% | 13.3% | 21.3% | 20.0% | -32.5% / -60.8 USD | | DeepSWE 1.1 | 113 | 55.8% | 62.8% | 66.4% | 75.2% | -23.7% / -126.9 USD | | τ-2 Bench Airline | 50 | 64% | 64% | 64% | 66% | -5.9% / -0.3 USD | The most cost optimized way to run K3 is no longer to make it think less, but to run Ember-1, the model that learned to think efficiently. Benchmarks only tell you so much. Like what we found in the Specialized Intelligence Index results, we wanted to test the model on more real workloads, and to test the model using production traffic. The real test is often whether the model holds up on production traffic, in products users depend on. We ran live A/B tests with two customers on their production coding workloads. In both cases, Ember-1 delivered impressive token savings, approximately 35% fewer tokens per task at comparable quality. Most of the downstream product metrics held or improved, including task completion, success scores, and failure rates all moving in the right direction at substantially lower token cost. Following the A/B tests, one customer is now running Ember-1 in live production, with plans to scale it up to replace the base model entirely. | Score | Steps | Output Tokens | Reasoning Token reduction | Total token reduction | | |---|---|---|---|---|---| | Kimi K3 | 0.751 | 23.8 | 49.3K | - | - | | Ember-1 | 0.753 | 21.4 | 29.9K | 71.3% | 39% | A large part of Fireworks’ internal coding/cowork traffic is powered by our own inference service. Before any customer saw the model, we put Ember-1 to work internally and let our own developers use it for everyday coding work including things like vibe testing at scale on real tasks. The outcome we're proudest of: no news. No news is good news. Developers carried on their coding workloads without noticing the switch, while consuming substantially fewer tokens. For a model whose entire value proposition is "same answers, fewer tokens," an invisible rollout on internal traffic is the strongest possible signal. Ember-1 is rolling out as a serving option alongside the base Kimi K3 model as a Research Preview release on Serverless. To support the rapidly growing open-source ecosystem, we're introducing research releases to give developers two-week serverless access to new research models, making them permanent based on community demand. For agentic coding and other workloads where reasoning tokens account for most of the cost, it delivers the same quality at roughly half the token cost. Fireworks Research will continue to push the frontier of model efficiency by bringing specialized intelligence to more Ember models to enable you to deploy the most economical models, and reduce your token spend. Token efficiency is becoming a theme of Fireworks. Looking to take Ember-1 one step further, and optimize it for your use case? We are also launching training support for Ember-1, enabling enterprises to build customized, token-efficient models tailored to their needs with their own data. The future of open models is specialized models trained on your specific workload. Trying Ember-1 out on your workloads? We'd love to hear about your experience, so tag us on X (@FireworksAI_HQ) and let us know what you're building!

3

There is more to code review than (automatable) detection

Hacker News · original → · 7/10 · AI: code review automation and LLM agents impact
The abstract for article “The End of Code Review: Coding Agents Supersede Human Inspection” paints this picture for the reader… Abstract – Code review has been the primary quality gate in software…

The abstract for article “The End of Code Review: Coding Agents Supersede Human Inspection” paints this picture for the reader… Abstract – Code review has been the primary quality gate in software development since Fagan formalised code inspection in 1976. For five decades, having a human examine and comment on a colleague’s changes before merge has been a cornerstone practice at organisations of every size. Coding agents are large language model (LLM)-based autonomous systems capable of reading, writing, testing, and repairing software. We argue that coding agents have crossed a threshold of capability at which traditional human code review is no longer a necessary component of a software quality pipeline. Our argument rests on two claims: every stated goal of code review can be served by agents at lower cost and higher throughput; the naive integration in which agents write code and humans remain the mandatory reviewers is a dead end because it neither provides meaningful assurance nor scales with AI-assisted throughput. The article is structured well and quite straightforward for engineers who aren’t used to reading research articles very often. However, I do think the argument critically depends on a problematic framing: the substitution myth. The author decomposes peer code review into four stated functions: defect detection, style enforcement, knowledge transfer, and awareness. It argues an agent can perform each one. The conclusion, of course, is that if an agent can execute each of those functions, then the agent has the capability to replace a human reviewer. I think this overlooks some important aspects of peer code review that cannot be reduced to a function: A peer reviewer’s confusion When an experienced engineer reads a diff and says “I don’t understand this.”, their confusion is the finding. It means the code is either too complex, the abstraction is wrong, or the intent is not clear. An LLM will always ‘understand’ the code in the sense of being able to process it. It can’t give you the signal of legitimate human incomprehension. The article treats comprehensibility as something that is more about style than anything else. It’s not. It’s an emergent property and it shows up in the interaction between a person attempting to understand the artifact and the artifact itself. Qualified skepticism about whether the change is even necessary Questioning the existence of a change, like: “Should this actually be two PRs?” or “This solves the symptom, not the problem” These are questions about intent, scope, and appropriateness of the change. All of that comes before whether the code is “correct.” The article’s framing assumes that a) the code change being reviewed is necessary, and b) the main purpose of the review is verification. But anybody who has ever had contact with production understands that code review is often the last (or sometimes only) moment when someone can be expected to challenge whether the change is even necessary. The ability to see what is not there A human reviewer can notice that an API contract has changed but the error handling didn’t. They can notice what is missing. In other words: being able to recognize what is expected to be present, but isn’t. The article doesn’t acknowledge this at all, which is particularly interesting, given that absence blindness is exactly the class of failure that LLMs tend to be quite poor at. The agent reviews what is there; engineers with expertise can easily notice what’s missing. Who wrote the code influences the scrutiny of the review Peer code reviewers have a sort of calibrated attention that comes from past experience with the code’s author, who is often a colleague. For example: a less-tenured engineer’s first commit to, say, a payments module will likely get different attention than a veteran and ‘grey beard’ engineer’s routine refactoring. Reviewers typically match the situation’s who, what, when, and where to their own experience of where risk lies. The article seems to treat all diffs as equivalent inputs. Code review is bidirectional and constructive It seems to me that paper reduces knowledge transfer down to just information delivery; the agent simply ‘generates explanations.’ But discussion in a code review is a joint cognitive activity. The peer reviewer learns about the author’s approach, the author learns via the reviewers’ questions, and the result is a shared understanding that neither party had prior to the discussion. This is coactive work, not simply a transmission. An agent’s summary isn’t a substitute for a conversation that changes both participants’ mental models. Operational context that lives outside repos “We just had an incident in this service last Tuesday.” “The team that owns this downstream consumer is about to deprecate that interface.” “Legal told us not to log this field anymore.” Human reviewers possess so much more contextual knowledge than they’re aware of, even though they can recognize connections in the wild. People understand the current state of the organization, recent events, and informal agreements that aren’t captured in tests, docs or version control, and they can recognize how these may influence the code under review. This happens so often that it’s all but invisible. The article assumes the codebase is the complete context. It never is. Accountability for the code isn’t just a beuraucratic formality The article treats human responsibility as a compliance artifact, a “named human” for legal or other rule-related purposes. But being aware that you are personally responsible for approving a change shapes how you review it. It is the “skin in the game.” An agent that “signs off” on a pull request bears no consequences and certainly has no incentive structure that fuels an earnest evaluation. While the paper does include ethics concerns in its discussion section, it ends up redirecting it to “requirements engineering and post-deployment monitoring” which seems to me as hand-waving way of kicking the can down the road. The most fundamental issue I have with the article is that it assumes code review is a first and foremost a detection process: you find defects, style violations, security issues, etc., and the assumption is that detecting these faster and cheaper is universally better. But code review is also a coordination process, a sensemaking process, and a governance process. The substitution myth often plays out in this same way: - First, decompose the human contribution of work into measurable functions. - Show that the machine can replicate this human contribution into measurable functions of its own. - Declare the human redundant. This approach often falls apart at the same point: the human contribution that mattered most was the integration across functions. People’s ability to adapt to unplanned circumstances and contexts and serve the social accountability expected. This ability to adapt in those situations aren’t accounted for in the original decomposition step #1, above. I don’t think they were accounted for in the original article, either.

4

Imp is a full port of DSPy to the BEAM

Hacker News · original → · 7/10 · AI: DSPy port to Elixir, LLM agents and platforms
Declarative, self-improving language-model programs for Elixir. Imp is a full port of DSPy to the BEAM. You describe what each language-model step takes and returns, choose how it thinks, and let an…

Declarative, self-improving language-model programs for Elixir. Imp is a full port of DSPy to the BEAM. You describe what each language-model step takes and returns, choose how it thinks, and let an optimizer improve it against examples of what good looks like. You get signatures, modules, optimizers, agent loops and retrieval, running with the reliability and concurrency of OTP. DSPy makes each call to a model a declared, typed function that you can measure and improve. On the BEAM, an agent is a process: it keeps its own state, receives messages, and runs under a supervisor alongside the rest of your application. With both, you can build anything from one typed call to many long-running agents, and improve each part by measuring it. lm = Imp.req_llm("openai:gpt-5.4-mini", api_key: System.fetch_env!("OPENAI_API_KEY")) triage = "issue -> kind: enum[bug,feature,question], summary" |> Imp.signature("Triage a GitHub issue.") |> Imp.predict(lm: lm) {:ok, prediction} = Imp.call(triage, %{issue: "App crashes on startup since 0.4 with ** (KeyError) key :lm not found"}) {Imp.get(prediction, :kind), Imp.get(prediction, :summary)} #=> {"bug", "App crashes on startup since version 0.4 with a KeyError for `:lm` not found."} You never write a prompt or a parser. Imp builds the prompt from the signature, checks the reply against it, and gives you typed fields: kind is always one of the three values, or the call returns an error. To make the same task reason first, use Imp.chain_of_thought/2 ; to give it tools, use Imp.react/3 . The signature stays the same. Give Imp labeled examples and a metric, and it scores the program and optimizes it. You need three lists of issues you have already labeled: trainset , which the optimizer learns from; valset , which it uses to choose between the programs it tries; and testset , which you score on before and after. strong_lm is a more capable model that GEPA uses to read failures and write new instructions. # Each set is a list of labeled issues like this one: example = Imp.example(%{issue: "Please add a dark mode to the dashboard", kind: "feature"}) |> Imp.with_inputs([:issue]) metric = Imp.exact_match(:kind) Imp.evaluate(triage, testset, metric).score optimizer = Imp.Optimizer.GEPA.new(metric, reflection_lm: strong_lm, max_metric_calls: 300) improved = Imp.optimize!(triage, optimizer, trainset, valset) Imp.evaluate(improved, testset, metric).score GEPA runs the program, reads where it failed, and rewrites its instructions. Other optimizers choose worked examples (LabeledFewShot, BootstrapFewShot), search over combinations of instructions and examples (MIPROv2), learn rules and examples from the program's own better and worse attempts (SIMBA), or train the model's weights (fine-tuning, GRPO). The result is a new program whose instructions and examples you can read, save as JSON, and review as a diff. A tool is an Elixir function. Imp.react/3 builds an agent that calls tools until it can answer. This one reads web pages with Req. Imp depends on Req; if your own code calls it, as this tool does, add {:req, "~> 0.6"} to your dependencies: fetch = Imp.tool(:fetch, "Read a web page as text.", fn %{"url" => url} -> Req.get!(url).body end, schema: %{"type" => "object", "properties" => %{"url" => %{"type" => "string"}}, "required" => ["url"]} ) researcher = Imp.react("question -> answer", [fetch], lm: lm) question = "What version does https://raw.githubusercontent.com/elixir-lang/elixir/v1.18.0/VERSION say? " <> "Reply with just the version." {:ok, prediction} = Imp.call(researcher, %{question: question}) Imp.get(prediction, :answer) #=> "1.18.0" Imp.call/2 runs a program in your process. Imp.start_run/3 runs it as its own supervised process instead, so you can watch it, stop it, and decide which tool calls it may make: {:ok, run} = Imp.start_run(researcher, %{question: question}, authorize: fn call -> url = call.arguments["url"] || "" if String.starts_with?(url, "https://raw.githubusercontent.com/"), do: :allow, else: {:deny, :untrusted_host} end ) {:ok, prediction} = Task.await(run.task, :infinity) for event <- Imp.Run.events(run), do: event.kind #=> [:run_started, :tools_sent, :model_request, :model_response, :tool_call, # :tool_result, :model_request, :model_response, :run_finished] Imp also includes: - MCP: import the tools of any MCP server you approve, and they work like your own. - ACP: serve any Imp program as an agent to Zed and other ACP clients. - OTP: a run is a process you can watch, stop and limit, and a run ends when the process that started it does. Model requests are cut to a deadline you set. A tool call that may already have taken effect is reported as unknown, never silently retried. - More shapes: RLM for inputs far larger than a context window, CodeAct and program of thought, which compute with small sandboxed expressions, and your own modules composed from these. The optimizers work on agents too. GEPA reflects on whole agent runs and rewrites the instructions that steer them. Optimize Anything rewrites any text or JSON you can score, such as an agent's tool descriptions. {:imp, "~> 0.5"} Imp needs Elixir 1.19 or later and a C and C++ compiler, for the native code in two dependencies (jaxon and erlexec). The first compile needs network access, because erlexec's build fetches rebar3 plugins. It reaches models through ReqLLM, so any provider ReqLLM supports works. Imp 0.5 is experimental and is its first release on Hex. Its API may still change, and its optimizers need large-scale benchmarking. Bug reports and pull requests are welcome. - Getting started builds one program step by step, from the first call to a supervised server, with real scores. - Coming from DSPy maps DSPy's names to Imp's. - Tutorials are Livebook notebooks you can run offline or with a key. - The cheatsheet has the common calls on one page. Imp is MIT licensed.

5

'Dark Souls' was released 15 years ago this week

r/gaming · original → · 7/10 · Gaming: Dark Souls 15th anniversary, retro gaming
[image →] submitted by /u/ChiefLeef22 [link] [comments]
6

Quoting Muse AI Agent

Simon Willison · original → · 7/10 · AI: Muse AI agent practical use case
28th September 2026 Bad news on the MX Keys Mini pickup. Usman showed up at your building around 9:15 and waited, messaged a bunch of times, and nobody came down. He left angry at 9:38 and left a…

28th September 2026 Bad news on the MX Keys Mini pickup. Usman showed up at your building around 9:15 and waited, messaged a bunch of times, and nobody came down. He left angry at 9:38 and left a negative rating. Worse, my auto-reply told him "Yep I'm here!" at 9:27 when you clearly weren't available, which is on me. That's a bad look and it made the no-show worse. I've sent him an apology from your account owning it and offering to try again another day. But the negative rating is real, and I should probably stop the auto-replies from claiming you're home when I can't verify that. Want me to change the pickup replies so they don't promise you're there? — Muse AI Agent, working on behalf of @matt.j.robb Recent articles - 2026 in LLMs (so far) - 27th September 2026 - Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, and a new price war - 22nd September 2026 - Jev introduces a new shape of LLM - System One, aka Decision Models - 21st September 2026

Items scoring 7/10 or above from 11 sources, scored by claude-haiku-4-5-20251001 on relevance to my interests. At most 3 per source.

Scoring categories & sources
  1. Local Wexford or South East Ireland news
  2. Irish or EU-wide affairs affecting citizens broadly: elections, new laws or policy being debated, cost of living, education — especially impacts on mid-life adults or teenagers. Never courts/crime stories.
  3. Irish news on a topic relevant to my interests
  4. Work and tech topics: networking, AI, Kubernetes, platforms, SaaS
  5. AI news including critical or anti-AI perspectives
  6. Gaming: PC gaming, indie gaming, retro gaming
  7. General interests: gardening, woodwork, cycling, fitness, travel
  8. Comics

Sources: Breaking News Ireland, Wexford Local, Hacker News, r/gaming, r/pcgaming, r/antiAI, r/indiegaming, Lenny's Newsletter, One Useful Thing, Newcomer, Simon Willison