Three counties in Ireland now require a single first-time buyer to earn a six-figure salary to buy a typical first home, according to a new study. That figure is up fromm just one county a year ago,…
Three counties in Ireland now require a single first-time buyer to earn a six-figure salary to buy a typical first home, according to a new study.
That figure is up fromm just one county a year ago, Dublin. A single buyer now needs €108,000 in Dublin, €101,250 in Wicklow and €100,238 in Kildare, according to new research from Chill Insurance.
The research also found that in every one of the 26 counties, a single buyer earning the national median salary of €44,816 would fall short of the income required to buy a typical first home.
The research describes a "country of two extremes".
At the lower end, even Longford, the most affordable county, requires a first-time buyer to earn €48,375 for a typical first home. At the top, Dublin, Wicklow and Kildare each now require a six-figure salary.
In all three, a single buyer on the local median income would need to more than double their salary to qualify for a mortgage on a typical first home in their county. In Dublin, the required salary is €108,000 against a local median income of €49,224, a shortfall of €58,776, the largest in the country.
The research compared the median first-time buyer home price in each county, using the Central Statistics Office's Residential Property Price Index, with median local incomes from CSO earnings data.
These figures were then assessed against the standard Central Bank lending rules, under which a buyer can borrow up to 90 per cent of a home's value and up to four times their gross income.
Economist Austin Hughes said: "With Irish house prices continuing to increase faster than wages, first-time buyers are finding it increasingly difficult to find a property they can afford, not just in Dublin or major urban centres but right across the country.
"The Chill report emphasises the nationwide nature of the problem would-be homebuyers face. Even allowing for ongoing wage growth, in no county is the typical first-time buyer property accessible at present to the median wage-earner, unless they are in a position to get additional financial support from the State and/or their family.
"With would-be buyers in Dublin and surrounding counties now needing around twice the typical wage in those areas, and approaching that in Cork and Galway, this 'purchase gap' is growing and spreading in a worrisome manner."
Hacker News
· original →
· 8/10
· AI: ChatGPT Work product analysis and capabilities
Understanding ChatGPT Work 30th August 2026 OpenAI announced ChatGPT Work on July 9th, and have been furiously iterating on it ever since. It is an extraordinarily confusing and very powerful…
Understanding ChatGPT Work
30th August 2026
OpenAI announced ChatGPT Work on July 9th, and have been furiously iterating on it ever since. It is an extraordinarily confusing and very powerful product. Here’s what I’ve figured out about it so far.
ChatGPT Work is actually two products
The more interesting version of ChatGPT Work is the one that runs in the cloud. This can be accessed via chatgpt.com or through the ChatGPT mobile apps. Let’s call it Work Cloud.
If you install the ChatGPT desktop app—the app that used to be called Codex—you gain access to a thing called ChatGPT Work that can access files and run programs directly on your computer. Let’s call that one Work Local. This one feels more like regular Codex re-skinned to be less intimidating to non-software-developers.
For the rest of this article I’m going to talk exclusively about Work Cloud.
Work is for paid subscribers only
Right now, ChatGPT Work (in both flavors) is available only to $20/month and up subscribers. Free users and $8/month Go users do not have access.
Work has features that aren’t available in Chat
The interface for accessing Work is a tab selector, which presents it as an alternative to Chat:
The obvious question is when should I use Chat, and when should I use Work?
OpenAI’s official answer to that question is:
Use Chat when you want an answer, explanation, brainstorm, or short draft. Use ChatGPT Work when you want ChatGPT to complete a task with a clear outcome, such as a brief, deck, analysis, recurring update, workflow, or file you can review and use.
I find that almost entirely useless, because I’ve been using regular ChatGPT Chat for all of those task categories for years!
The better question then is what features does Work have that are missing from Chat?
After extensive experimentation I think I’ve mostly figured that out:
- Options to use Luna and Terra in place of Sol
- A code execution environment with Internet access
- A headless Chrome browser
- A persistent filesystem shared between sessions
- The ability to publish ChatGPT Sites
- The ability to run sub-agent sessions with Sol, Luna, and Terra
- Scheduled prompt automations (may be in ChatGPT Chat too)
Model selection
In Work, you get the option to pick GPT-5.6 Sol, Luna, or Terra, each with Light, Medium, High, Extra High, Max, or Ultra reasoning levels. You can also pick GPT-5.5 at Light, Medium, High, or Extra High.
These look to be the same models that are available through the OpenAI API.
Chat offers a different selection: 5.6 Instant, Medium, High, Extra High, and Pro (actually Extra High and Pro are only available for $100/month+ subscribers—$20/month subscribers cap out at High). It doesn’t explain if those are Luna or Terra or Sol (I’m assuming Sol?). 5.6 Pro appears to be exclusive to Chat, with no equivalent in Work.
My current understanding from using Codex is that Ultra is a special mode that more eagerly delegates to sub-agents.
I believe ChatGPT Work sessions are billed against your Codex allowance, while ChatGPT Chat Sessions get their own, separate allowance. This may help explain the model availability differences.
Code execution with Internet access!
As a long-time fan of the Code Interpreter pattern—pioneered by OpenAI in 2023—this is by far the most exciting feature of ChatGPT Work (Cloud) for me.
The code execution environment can now talk to the rest of the internet!
ChatGPT Chat can’t do this—if you ask it to install additional software packages or interact with websites or APIs that access will be blocked by the container proxy.
(Weirdly, back in January it grew the ability to install packages, but that doesn’t seem to work any more. I wish they had better changelogs!)
Claude’s equivalent container has allowed restricted internet access since it launched last September. Claude can install packages from PYPI and NPM and clone repositories from GitHub. But that is about it: the allowlist of domains is very short.
ChatGPT Work allows a whole lot more than that. It can be configured with a specific list of allowed domains, but the default appears to be open to all.
This makes Work an incredibly useful tool. You can have it clone GitHub repositories, install their dependencies, then use them to interact with the rest of the web!
A full, headless Chrome browser
Another killer feature of ChatGPT Work is the browser tool. ChatGPT Work can launch a full Chrome instance, load websites, fill out forms, and take screenshots.
If a site requires sign in the browser can prompt you to take over and enter both passwords and 2FA codes, without round-tripping those credentials through the model itself.
It can even run JavaScript against the DOM of loaded pages. I prompted:
Load simonwillison.net in your browser and extract the headings using JavaScript
ChatGPT Work fired up a browser instance and ran the code:
await tab.playwright.evaluate(() => {
return Array.from(document.querySelectorAll("h1,h2,h3,h4,h5,h6"), heading => ({
level: heading.tagName.toLowerCase(),
text: heading.innerText.trim().replace(/\s+/g, " "),
id: heading.id || null
}));
});
This feels a lot like my shot-scraper javascript tool, only now I can access it on my phone!
A persistent, shared filesystem
ChatGPT Chat gets a fresh filesystem for each chat session. These cannot be accessed from any other session.
In ChatGPT Work each session gets its own scratch folder—named something like /workspace/scratch/e00a0a017944
—but each of those are persisted across sessions, so you can access files from previous chats. I have 171 folders in /workspace/scratch
right now!
As far as I can tell that /workspace
volume is mounted to all Work sessions that are currently running—file edits from one can be instantly seen by the others. They don’t seem to share the same process space though, and localhost servers running in one can’t be accessed from another.
ChatGPT Sites
ChatGPT Work has the ability to build and deploy entire websites, using Cloudflare Workers. These can have HTML and JavaScript and can run server-side features too, including stateful features on top of Cloudflare D1 and R2.
Here’s a simple site I built with this feature:
london-pelicans-in-her-piety.simonw.chatgpt.site
My prompt was:
Figure out all of the places in London with a pelican in her piety, then turn that into a JSON file and build a ChatGPT sites site about them
(A pelican in her piety is a fascinating piece of medieval Christian imagery—once you know about them you’ll find them all over the place.)
These sites default to being private to the user that created them, but you can make them public and (on team plans) share them with other specific individuals.
Sub-agents with Sol, Luna, and Terra
There’s not much to say about this one. ChatGPT Chat can’t run sub-agents. ChatGPT Work can. This is very much a power-user feature: if you are running a complex project that can benefit from multiple parallel agents working together, Work can do that.
Scheduled prompt automations
Another feature that seems to have migrated from regular ChatGPT to ChatGPT Work at some point. You can prompt ChatGPT Work like this:
run a search to see if Waymo have announced a launch date for Half Moon Bay every day at 8am
This will schedule a prompt to run on that frequency. These prompts can decide that nothing interesting has happened, or they can decide to notify you of some new information.
Update: Actually this seems to work in ChatGPT Chat as well.
It’s still worth noting here though, as it can be used in conjunction with other ChatGPT Work exclusive features. You can set a scheduled task to update a ChatGPT Site on an hourly basis, for example.
Is this safe?
An open question for me right now is how safe all of this stuff is.
My lethal trifecta model warns about the risks inherent in any agent system that combines access to private data with exposure to untrusted content and a way to communicate stolen information back to an attacker.
ChatGPT Work combines all three!
I’d love to hear more from OpenAI about how they protect ChatGPT Work sessions against prompt injection attacks. I expect their answer is the same auto-review mechanism as Codex.
OpenAI could make this a lot less confusing
Figuring this all out took way more work than it should have.
I think there are two key problems here:
- OpenAI explain Work in terms of what it’s for, not what it actually does
- OpenAI still insist on hiding their system prompts and tools descriptions
If the ChatGPT Work documentation included the exact system prompt and tool descriptions used by the agent I wouldn’t have needed to write this post.
A list of all the tools
Shortly after publishing this article I had an idea. I started a fresh Work session and prompted:
Build a site that lists every one of your tools - nearly grouped into categories - and for each one explain what it does. Try to exactly duplicate arguments and tool descriptions where possible. Design aesthetic should be technical docs, minimal flare
Here’s the site it built, which includes details of 223 registered tools—though 6 of those are from my own personal MCPs served via datasette-mcp.
And a whole lot of Skills
I noticed that the only browser-related tool in the list was web.run, which has methods for running searches, opening URLs, and clicking links, but didn’t look like the full story in regards to headless browser automation.
This made me suspicious that something was missing, so I told the ChatGPT Work session that built that tools reference site:
Add full copies of every skill to the website (separate pages linked to from the homepage)
It turns out ChatGPT Work uses a lot of skills—44 in fact!
The control-browser skill explains how the browser works:
Run browser setup code through the Node REPL
js
tool. In this environment the callable tool id typically appears asmcp__node_repl__js
. [...]The ability to interact directly with the browser is exposed through the
browser-client
runtime via theagent.browsers.*
API. Before trying to interact with it, you MUST emit and read the complete documentation returned byawait browser.documentation()
in one go.
So I told Work:
Add the full output of await browser.documentation() to the bottom of the /skills/control-browser page
And now you can read that on /skills/control-browser as well.
A few more interesting Skills:
-
documents for creating
.docx
files -
imagegen with tips on creating images with the
image_gen
tool - pdf for both reading and rendering PDFs
-
Spreadsheets for manipulating
.xlsx
,.xls
,.csv
,.tsv
- sites:sites-building for creating ChatGPT Sites
- openai-docs for answering questions about itself
- data-analytics:build-dashboard for building data dashboards
More recent articles
- Conceptual integrity and counting lines of code - 19th August 2026
- Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things - 16th August 2026
Lenny's Newsletter
· original →
· 8/10
· AI/platforms: OpenAI product lead on persistent AI coworkers
[image →]Tara Seshan leads product for Codex and ChatGPT Work at OpenAI (alongside previous podcast guest Andrew Ambrosino, who’s her engineering manager). Before OpenAI, Tara spent over six years…
Tara Seshan leads product for Codex and ChatGPT Work at OpenAI (alongside previous podcast guest Andrew Ambrosino, who’s her engineering manager). Before OpenAI, Tara spent over six years at Stripe, where she joined as one of the first five product managers. She went on to lead product for Watershed, which Time magazine named one of the best inventions of 2022, and she is also a founder and Thiel Fellow. Most personally meaningful to me: Tara is one of the three inaugural Lenny’s Newsletter Fellows, a program I ran a couple of years ago to spotlight the most exciting up-and-coming product leaders.
Simon Willison
· original →
· 8/10
· AI: ChatGPT Work product deep-dive and analysis
Understanding ChatGPT Work 30th August 2026 OpenAI announced ChatGPT Work on July 9th, and have been furiously iterating on it ever since. It is an extraordinarily confusing and very powerful…
Understanding ChatGPT Work
30th August 2026
OpenAI announced ChatGPT Work on July 9th, and have been furiously iterating on it ever since. It is an extraordinarily confusing and very powerful product. Here’s what I’ve figured out about it so far.
ChatGPT Work is actually two products
The more interesting version of ChatGPT Work is the one that runs in the cloud. This can be accessed via chatgpt.com or through the ChatGPT mobile apps. Let’s call it Work Cloud.
If you install the ChatGPT desktop app—the app that used to be called Codex—you gain access to a thing called ChatGPT Work that can access files and run programs directly on your computer. Let’s call that one Work Local. This one feels more like regular Codex re-skinned to be less intimidating to non-software-developers.
For the rest of this article I’m going to talk exclusively about Work Cloud.
Work is for paid subscribers only
Right now, ChatGPT Work (in both flavors) is available only to $20/month and up subscribers. Free users and $8/month Go users do not have access.
Work has features that aren’t available in Chat
The interface for accessing Work is a tab selector, which presents it as an alternative to Chat:
The obvious question is when should I use Chat, and when should I use Work?
OpenAI’s official answer to that question is:
Use Chat when you want an answer, explanation, brainstorm, or short draft. Use ChatGPT Work when you want ChatGPT to complete a task with a clear outcome, such as a brief, deck, analysis, recurring update, workflow, or file you can review and use.
I find that almost entirely useless, because I’ve been using regular ChatGPT Chat for all of those task categories for years!
The better question then is what features does Work have that are missing from Chat?
After extensive experimentation I think I’ve mostly figured that out:
- Options to use Luna and Terra in place of Sol
- A code execution environment with Internet access
- A headless Chrome browser
- A persistent filesystem shared between sessions
- The ability to publish ChatGPT Sites
- The ability to run sub-agent sessions with Sol, Luna, and Terra
- Scheduled prompt automations (may be in ChatGPT Chat too)
Model selection
In Work, you get the option to pick GPT-5.6 Sol, Luna, or Terra, each with Light, Medium, High, Extra High, Max, or Ultra reasoning levels. You can also pick GPT-5.5 at Light, Medium, High, or Extra High.
These look to be the same models that are available through the OpenAI API.
Chat offers a different selection: 5.6 Instant, Medium, High, Extra High, and Pro (actually Extra High and Pro are only available for $100/month+ subscribers—$20/month subscribers cap out at High). It doesn’t explain if those are Luna or Terra or Sol (I’m assuming Sol?). 5.6 Pro appears to be exclusive to Chat, with no equivalent in Work.
My current understanding from using Codex is that Ultra is a special mode that more eagerly delegates to sub-agents.
I believe ChatGPT Work sessions are billed against your Codex allowance, while ChatGPT Chat Sessions get their own, separate allowance. This may help explain the model availability differences.
Code execution with Internet access!
As a long-time fan of the Code Interpreter pattern—pioneered by OpenAI in 2023—this is by far the most exciting feature of ChatGPT Work (Cloud) for me.
The code execution environment can now talk to the rest of the internet!
ChatGPT Chat can’t do this—if you ask it to install additional software packages or interact with websites or APIs that access will be blocked by the container proxy.
(Weirdly, back in January it grew the ability to install packages, but that doesn’t seem to work any more. I wish they had better changelogs!)
Claude’s equivalent container has allowed restricted internet access since it launched last September. Claude can install packages from PYPI and NPM and clone repositories from GitHub. But that is about it: the allowlist of domains is very short.
ChatGPT Work allows a whole lot more than that. It can be configured with a specific list of allowed domains, but the default appears to be open to all.
This makes Work an incredibly useful tool. You can have it clone GitHub repositories, install their dependencies, then use them to interact with the rest of the web!
A full, headless Chrome browser
Another killer feature of ChatGPT Work is the browser tool. ChatGPT Work can launch a full Chrome instance, load websites, fill out forms, and take screenshots.
If a site requires sign in the browser can prompt you to take over and enter both passwords and 2FA codes, without round-tripping those credentials through the model itself.
It can even run JavaScript against the DOM of loaded pages. I prompted:
Load simonwillison.net in your browser and extract the headings using JavaScript
ChatGPT Work fired up a browser instance and ran the code:
await tab.playwright.evaluate(() => {
return Array.from(document.querySelectorAll("h1,h2,h3,h4,h5,h6"), heading => ({
level: heading.tagName.toLowerCase(),
text: heading.innerText.trim().replace(/\s+/g, " "),
id: heading.id || null
}));
});
This feels a lot like my shot-scraper javascript tool, only now I can access it on my phone!
A persistent, shared filesystem
ChatGPT Chat gets a fresh filesystem for each chat session. These cannot be accessed from any other session.
In ChatGPT Work each session gets its own scratch folder—named something like /workspace/scratch/e00a0a017944
—but each of those are persisted across sessions, so you can access files from previous chats. I have 171 folders in /workspace/scratch
right now!
As far as I can tell that /workspace
volume is mounted to all Work sessions that are currently running—file edits from one can be instantly seen by the others. They don’t seem to share the same process space though, and localhost servers running in one can’t be accessed from another.
ChatGPT Sites
ChatGPT Work has the ability to build and deploy entire websites, using Cloudflare Workers. These can have HTML and JavaScript and can run server-side features too, including stateful features on top of Cloudflare D1 and R2.
Here’s a simple site I built with this feature:
london-pelicans-in-her-piety.simonw.chatgpt.site
My prompt was:
Figure out all of the places in London with a pelican in her piety, then turn that into a JSON file and build a ChatGPT sites site about them
(A pelican in her piety is a fascinating piece of medieval Christian imagery—once you know about them you’ll find them all over the place.)
These sites default to being private to the user that created them, but you can make them public and (on team plans) share them with other specific individuals.
Sub-agents with Sol, Luna, and Terra
There’s not much to say about this one. ChatGPT Chat can’t run sub-agents. ChatGPT Work can. This is very much a power-user feature: if you are running a complex project that can benefit from multiple parallel agents working together, Work can do that.
Scheduled prompt automations
Another feature that seems to have migrated from regular ChatGPT to ChatGPT Work at some point. You can prompt ChatGPT Work like this:
run a search to see if Waymo have announced a launch date for Half Moon Bay every day at 8am
This will schedule a prompt to run on that frequency. These prompts can decide that nothing interesting has happened, or they can decide to notify you of some new information.
Update: Actually this seems to work in ChatGPT Chat as well.
It’s still worth noting here though, as it can be used in conjunction with other ChatGPT Work exclusive features. You can set a scheduled task to update a ChatGPT Site on an hourly basis, for example.
Is this safe?
An open question for me right now is how safe all of this stuff is.
My lethal trifecta model warns about the risks inherent in any agent system that combines access to private data with exposure to untrusted content and a way to communicate stolen information back to an attacker.
ChatGPT Work combines all three!
I’d love to hear more from OpenAI about how they protect ChatGPT Work sessions against prompt injection attacks. I expect their answer is the same auto-review mechanism as Codex.
OpenAI could make this a lot less confusing
Figuring this all out took way more work than it should have.
I think there are two key problems here:
- OpenAI explain Work in terms of what it’s for, not what it actually does
- OpenAI still insist on hiding their system prompts and tools descriptions
If the ChatGPT Work documentation included the exact system prompt and tool descriptions used by the agent I wouldn’t have needed to write this post.
A list of all the tools
Shortly after publishing this article I had an idea. I started a fresh Work session and prompted:
Build a site that lists every one of your tools - nearly grouped into categories - and for each one explain what it does. Try to exactly duplicate arguments and tool descriptions where possible. Design aesthetic should be technical docs, minimal flare
Here’s the site it built, which includes details of 223 registered tools—though 6 of those are from my own personal MCPs served via datasette-mcp.
And a whole lot of Skills
I noticed that the only browser-related tool in the list was web.run, which has methods for running searches, opening URLs, and clicking links, but didn’t look like the full story in regards to headless browser automation.
This made me suspicious that something was missing, so I told the ChatGPT Work session that built that tools reference site:
Add full copies of every skill to the website (separate pages linked to from the homepage)
It turns out ChatGPT Work uses a lot of skills—44 in fact!
The control-browser skill explains how the browser works:
Run browser setup code through the Node REPL
js
tool. In this environment the callable tool id typically appears asmcp__node_repl__js
. [...]The ability to interact directly with the browser is exposed through the
browser-client
runtime via theagent.browsers.*
API. Before trying to interact with it, you MUST emit and read the complete documentation returned byawait browser.documentation()
in one go.
So I told Work:
Add the full output of await browser.documentation() to the bottom of the /skills/control-browser page
And now you can read that on /skills/control-browser as well.
A few more interesting Skills:
-
documents for creating
.docx
files -
imagegen with tips on creating images with the
image_gen
tool - pdf for both reading and rendering PDFs
-
Spreadsheets for manipulating
.xlsx
,.xls
,.csv
,.tsv
- sites:sites-building for creating ChatGPT Sites
- openai-docs for answering questions about itself
- data-analytics:build-dashboard for building data dashboards
More recent articles
- Conceptual integrity and counting lines of code - 19th August 2026
- Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things - 16th August 2026
Changes to the income thresholds for the National Childcare Scheme will see thousands of families receiving more subsidies. From today, the income thresholds used to calculate Income Assessed…
Changes to the income thresholds for the National Childcare Scheme will see thousands of families receiving more subsidies.
From today, the income thresholds used to calculate Income Assessed subsidies will increase along with the Multiple Child Discount.
The changes will see the families of 47,000 children receive an increase in their childcare subsidies.
The National Childcare Scheme provides two types of financial support for families.
All eligible families may apply for a universal subsidy, regardless of family income, which is worth up to €96.30 per week, for a maximum of 45 weekly hours.
Alternatively, families may qualify for a higher income-assessed subsidy which varies depending on family circumstances, including income level and number of children under the age of 15.
The income thresholds used to calculate the income-assessed subsidies will increase.
As part of the changes, the lower income threshold has increased from €26,000 to €34,000, so families now earning €34,000 or less will be able to qualify for the maximum subsidy rate.
A family with a reckonable income of €34,000 will now receive a rate of €5.10 per hour for a child aged between 24-52 weeks. Prior to these changes, that same family would have received €4.40 per hour.
The upper income threshold will increase from €60,000 to €68,000, meaning that families now earning up to €68,000 will be able to avail of a higher income assessed rate.
By increasing the income-assessed thresholds and Multiple Child Discount, more families will benefit from higher levels of financial support under the National Childcare Scheme.
For example, a family with a reckonable income of €60,000 will now receive a rate of €2.84 per hour for a child aged between 24-52 weeks, compared to €2.14 per hour previously.
As a result of these changes, all existing income assessed subsidy recipients with an income between €26,000 and €60,000 will see an increase in their subsidy rate.
The Multiple Child Discount will also increase, meaning that a lower income is used to assess subsidy entitlements, enabling many families to receive a higher hourly subsidy rate.
The Multiple Child Discount will increase from €4,300 to €5,500 for families with two children and from €8,600 to €11,000 for families with three or more children.
Minister for Children Norma Foley said: “These increases in childcare subsidies under the National Childcare Scheme represent another important step forwards in making early learning and childcare more affordable for families.
“By increasing the income-assessed thresholds and Multiple Child Discount, more families will benefit from higher levels of financial support under the National Childcare Scheme.”
In total, more than €528 million will be paid to families this year under the National Childcare Scheme, out of total state investment of almost €1.5 billion in the early learning and childcare sector.
Hacker News
· original →
· 7/10
· AI: critical perspective on AI crawlers and system load
Creepy crawlies You've probably heard me complain about the “AI crawlers” before, but now I actually have some hard numbers I can put up to show their impact. In a few words, it's bad enough to…
Creepy crawlies
You've probably heard me complain about the “AI crawlers” before, but now I actually have some hard numbers I can put up to show their impact. In a few words, it's bad enough to create a constant “background radiation” of system load, permanently tying up a chunk of capacity spent on producing output that is only useful for a single purpose — feeding a learning model.
TL;DR: we spend more CPU cycles rendering commits for scrapers than we spend on all other kinds of legitimate access, including git clones. At any one time, across 5 geo-distributed nodes, there are 14 CPU cores doing nothing but rendering git commits as html.
Why is git.kernel.org “interesting” to crawlers
Linux development happens in the open — from git repositories you can clone, to discussion archives you can follow in real time. To a large language model, this is a goldmine of learning data, because all of this is not only immediately available, but is easy to filter in order to guarantee pure unadulterated pre-AI content. Training an LLM on content produced by the LLM gives it the equivalent of a digital prion disease, so when a source is guaranteed to be LLM-free, like the entire history of kernel commits, it's worth its weight in gold as a source of training data.
The stupidest way of doing it
We make almost everything clonable, because hey — we may not be around forever, so here — clone the repos. Also, clone the archives. Grab a copy just so we're not the only ones who have it all. Seriously, it's just a “git clone” away — and then you'll have the whole history.
For example, did you know you can clone the entirety of LKML and then do whatever you want with it? It's just git repos all the way down.
So, you'd think that something that pretends to be “Artificial Intelligence” would use the most efficient way of using our data for training purposes, right? Clone the repos, walk every commit. Done.
But no, let's in fact choose the stupidest possible way of doing it — by rendering everything as HTML commit by commit and then parsing it.
At the time of writing, linux.git is about 1.48 million commits. Oh, and we have about 922 forks of it on git.kernel.org — but don't worry, it's actually extremely efficient on the backend, since it's mostly the same objects in every fork.
Unless, of course, you're a scraper, in which case you have, oh, several BILLION valid URLs you can scrape, only to get 922 duplicates of the same 1.48 million commits — which is exactly what the scrapers are doing.
But wait, it's not just commits itself. You can also ask for patches, plain renders, diffs between arbitrary commits — cgit is happy to let you, which was perfect for the times when the Internet was for humans or crawlers who obeyed robots.txt, and is AWFUL right about now, because we can generate 1.2 METRIC BAJILLION valid URLs just for a single fork of linux.git.
Block them
Initially, this was the solution — look through the logs, find out which IPs are obvious scraper bots, and fail2ban them. At first, this was easy, because the bots helpfully told you who they were via their user-agent. Then, they wised up and started pretending that they were random vanilla browsers.
So, we started banning them by IP — after all, it's easy to figure out that an IP that is trying to grab every possible commit in a 8-year-old abandoned fork of linux is not really some lone Chrome on Windows user who is just furiously clicking every link that comes across their screen.
The bots then started fanning out to entire subnets, but this was still meh, because obviously an IP coming from Google Compute is just pretending to be a Firefox user. Banning the whole ASN was justified, even if this occasionally caught a random legitimate instance trying to automate link checking in commits.
Enter... your TV?
And... that's when things turned really, really ugly. Suddenly, the crawlers were coming from millions of random residential or mobile IPs, all pretending to be random modern browsers. An IP like that would make 4-5 requests and then never show up in the logs again. There was no point in banning them, because by the time you figured out that they were bots, they were already done with you. You just needlessly ballooned your firewall ruleset by adding IPs that would never be back.
They descended like swarms of locust, hit hard and fast until the system fell over and then moved on to the next target until you recovered. Then, they returned. Rinse. Repeat.
They still do that — welcome to the wonderful world of “proxy SDK monetization.” It's big business, and your TV is probably doing it.
Make them pay
When this first became a problem, oh, about a year ago, we naively thought that there was a way to make it stop. Just make the bots perform a task that would flip the economy of the whole thing upside-down by making them burn some cycles doing throwaway math. Like, calculate what string, when combined with their own IP and a secret we provide, would generate a sha256 sum with 4 leading zeroes.
In other words, we put Anubis in front of everything.
It was immediately extremely effective — the bots just gave up. For a few months, it was bliss: bots were blocked at the perimeter and gave up, moving on to easier targets; the users were mildly annoyed but tolerated it, and the Anubis stack was easy enough to deploy everywhere.
A few months later, the bots were back, solving difficulty 4. No problem, we said, let's raise difficulty to 5.
The legitimate users were more annoyed now. Difficulty 5 takes a few seconds to solve on a mobile device, and the phone gets uncomfortably warm as it's doing the number crunching. However, it was effective and bought us a few more months of peace.
Then... the bots started solving difficulty 5.
Where we are now
Today, git.kernel.org receives about 6M daily requests demanding to see random commits. Of these, 66% are still immediately batted away with the Anubis challenge, but 33% are now solving the math and getting through to the main site — because apparently what we have to offer is worth spending a ton of cycles to calculate the Anubis challenge.
It's impossible to tell with certainty which of these are bots and which are real humans — but chances are, if it's asking for an old commit in a random old fork, it's probably not a real developer trying to do their work.
With a bunch of generous assumptions, legitimate requests are only about 2% of git.kernel.org traffic — everything else are scrapers.
How bad is it?
At this point, we're not quite overwhelmed — if you visit git.kernel.org, it will likely be snappy and responsive. The thing that usually takes us down are not scraper bots, but poorly designed CI systems that try to do something stupid like shallow-clone stable.git from 20 different nodes, all at the same time. (Shallow clones are awful. Run your own damn mirror if you're going to do something nasty like that.)
However, you should know that out of the total of the 90 cores across 5 geo-distributed nodes, there are 14-16 cores that are constantly doing nothing but rendering commits for scrapers. On average, that's 20% of our entire capacity — except the swarms descend in waves and the actual graph is a lot more spiky than a 20% flatline.
Where does that leave us?
Unclear. Maybe the AI bubble bursts and we suddenly have a lot fewer entities out there trying to train their models. Alternatively, maybe they smarten up and stop consuming our data in the dumbest way possible.
In terms of what we're doing, we're turning off features to reduce the number of crawlable URLs and to gate off actions that are expensive for us to run. Expect to lose some functionality, at least when accessing our resources anonymously. Trust me, we hate it just as much as you, but at this point it's a necessity.
Worst of all, there are no simple solutions to the problem. Companies offering custom “AI” models still pop up daily, all of them hungry for training data. App makers are still looking for ways to turn a profit, so they will continue to turn your household appliances into attack vectors.
That said, we promise to still offer all of our data for download to anyone who asks. You just may have to jump through more hoops to get it.
Sorry. (Oblig. Canadian thing to say.)
Hacker News
· original →
· 7/10
· AI: diffusion language models technical deep-dive
An introduction to diffusion language models and the research advances that underlie today's diffusion LLMs. We describe the building blocks of recent open-source models, starting from simple…
An introduction to diffusion language models and the research advances that underlie today's diffusion LLMs. We describe the building blocks of recent open-source models, starting from simple masking diffusion, and including techniques for iterative refinement, post-training, and variable-length generation. Material is adapted from workshop talks and lectures at ICLR 2026 and MLSS 2026.
Two families of generative AI algorithms are widely used today. For continuous data such as images or video, the state-of-the-art approach is based on diffusion models. For discrete data such as text or code, the standard approach is instead autoregressive models. This article explores an alternative for discrete data, one built on the modern paradigm of diffusion.
Mainstream language models are autoregressive: they generate tokens left-to-right, one at a time, each conditioned on the tokens before it. This approach is powerful, but it also has inherent limitations:
Diffusion models take a different approach. Rather than producing text one token at a time, they generate the whole sequence at once, starting from an initial guess and iteratively refining it over a number of steps. This unlocks several advantages: generation can trade off speed and quality by using fewer or more steps, mistakes can be corrected along the way, and every step attends to bidirectional context.
Applying diffusion to language had long been an open problem. In 2024 the field reached a turning point, as diffusion models became competitive with autoregressive models on quality. By 2026, diffusion LLMs are a reality, with releases from leading industry labs — Mercury 2 (Inception Labs)
Before introducing diffusion for language, we start with a brief overview of Gaussian diffusion for image generation. We will then build up discrete diffusion by analogy.
The central concept underlying diffusion models is denoising. Instead of painting an image in one shot, a diffusion model produces images step by step, starting from pure random noise and removing a little of it at every step until a coherent image emerges. Generating an image through many small steps turns out to be far simpler than producing it all at once, and this is what makes diffusion models so effective.
How does a model learn to denoise? The trick is to teach it by showing examples of noise being gradually transformed into an image. Diffusion achieves this via two complementary processes. First, a forward process takes a clean source image and turns it into pure noise, one step at a time. Second, a reverse process learns to invert this transformation, turning pure noise back into an image; it is trained on the image-to-noise trajectories produced by the forward process.
The forward process takes a clean training image and produces a sequence of increasingly noisy images that trace a path from clean data to pure noise. It does this by mixing in a growing amount of random Gaussian noise at each step, until the image dissolves into pure static. This step requires no learning at all — we are simply adding noise — yet it is enormously useful, because it manufactures an endless supply of training data: examples of images being transformed into noise, and vice versa.
The reverse process is where the actual learning happens. We train a model to transform noise into images by following the steps produced by the forward process in reverse.
Concretely, given a noisy image, we train a machine learning model to separate the noise from the underlying image or, equivalently, to predict either the noise that was added or the clean image itself, since given the noisy input, knowing one determines the other. Once the model can do this, generation is simple: start from pure noise, ask the model to estimate and strip away a bit of it, and repeat. Each pass nudges the sample a little closer to something that looks like real data, until a clean image remains.
This forward/reverse recipe — corrupt data with noise, then learn to reverse the corruption one step at a time — is the blueprint for every diffusion model.
The main obstacle in bringing diffusion to language is deciding what "noise" should mean for discrete tokens. For example, the noise used in classical diffusion is Gaussian, and adding continuous Gaussian noise to categorical variables is not well-defined. Below we introduce one simple yet effective approach that defines noise via masking. Our group popularized this approach
The easiest way to understand masked diffusion is as an unmasking transformer. We train the model by taking clean sequences, masking a random fraction of their tokens, and asking a bidirectional transformer to fill in the blanks. If you know BERT, this is essentially BERT with a randomized masking rate — but unlike BERT, the resulting model is generative. You can think of masked diffusion as a generative BERT.
Once we trained the unmasking transformer, we can generate text by starting from a fully masked sequence and repeating two steps many times:
Each round leaves fewer positions masked, until the sequence converges to a clean sample from the model. Generation thus amounts to starting from a sequence full of blanks and gradually filling in words in an arbitrary order.
We can also understand a bit better why this process works by framing it as an analog of the Gaussian diffusion model we saw earlier. Just like Gaussian diffusion, masked diffusion can be described as a model consisting of a forward and a reverse process.
The goal of the forward process is to generate training data for the reverse process. Its output is a trajectory that starts from a datapoint and ends at a sequence of pure noise; the reverse process will then be trained to produce this trajectory in reverse.
The key challenge is deciding what "noisy" should mean. In Gaussian diffusion, we added varying amounts of white noise to an image. In masked diffusion, we instead randomly mask a fraction of the tokens in a discrete sequence. The amount of masking is governed by a schedule $\alpha_t$ — the probability that a given token remains unmasked — which plays the role of the signal-to-noise ratio in Gaussian diffusion. It starts at $1$ when $t = 0$ (a clean sequence) and decreases to $0$ when $t = 1$ (a fully masked sequence). The time variable $t$ indexes a path from clean to noisy data, and at time $t$ a partially masked sequence $z_t$ has, in expectation, a fraction $\alpha_t$ of its tokens unmasked.
We implement this process as a Markov chain over a sequence of variables $z_t$ indexed by $t$, with $z_0$ being the clean, unmasked sequence. For $s < t$, the chain defines $q(z_t \mid z_s)$ by masking each still-unmasked token of $z_s$ with probability $(\alpha_s - \alpha_t)/\alpha_s$. Running this Markov chain for a number of steps produces a trajectory going from clean data to fully masked noise.
Next, as in Gaussian diffusion, we train the reverse process to walk the sequence of increasingly masked latents in reverse — starting from a fully masked sequence and ultimately generating outputs similar to clean data.
Using Bayes' rule, we can derive the mathematically optimal reverse process $q(z_s \mid z_t, x)$ when the clean sequence $x$ is known
In practice, the final output $x$ is obviously unknown when we generate it. We therefore train a model $x_\theta(z_t)$ to predict the final clean sequence given the current state $z_t$ and apply the ideal reverse process $q(z_s \mid z_t, x)$ using the estimate $x_\theta(z_t)$ in place of the real $x$. More formally, we define the reverse process as a probability $p(z_s \mid z_t) = q\big(z_s \mid z_t, x_\theta(z_t)\big)$. This definition recovers the sampling algorithm we described earlier: at each step, we use the model $x_\theta(z_t)$ to fill in the blanks of $z_t$, and we keep a subset of these filled-in tokens in $z_s$.
Putting these pieces together gives us the mathematical definition of a masked diffusion language model (MDLM). The forward process $q(z_t \mid z_s)$ produces a trajectory from clean to fully masked data, and the reverse process $p(z_s \mid z_t)$ learns to undo it. Moreover, the reverse process defines a latent variable model $p(x, z_1, \dots, z_T)$ in which $T$ intermediate partially masked samples $z_1,...,z_T$ are latent variables. Generating from the reverse process $p(z_s \mid z_t)$ is the same as performing ancestral sampling from this model.
We can also look at the likelihood $\log p(x)$ of the model $p$ to assess its quality. In latent variable models this is intractable, so we resort to approximations via variational inference. For a masked diffusion language model, the evidence lower bound (ELBO) used to approximate the likelihood has a surprisingly simple form (assuming for simplicity $\alpha_t = 1-t$)
Let's unpack this formula. The inner term $\log p_\theta(x \mid z_t)$ is the likelihood of a clean sequence $x$ given a partially masked sequence $z_t$ sampled from the forward process. In other words, it is the cross-entropy loss between the predictions of our unmasking transformer and the true tokens — this is exactly the BERT loss!
Differently from BERT, this loss is averaged over all $t$, and hence over all possible masking rates, rather than a single fixed one. It is also normalized by $t$, the expected fraction of tokens that are masked (since $\alpha_t = 1-t$); this factor ensures that each BERT loss is normalized for the number of tokens over which the loss is taken.
In summary, MDLM is very similar to BERT, with two key differences:
Most interestingly, the evidence lower bound enables a principled comparison between autoregressive and diffusion language models using log-likelihood (or, equivalently, perplexity) — the standard metric for evaluating language models. While for a long time there was a substantial gap in perplexity between diffusion and autoregressive language models, simplified masked diffusion models were among the first to close much of this gap
As defined above, masked diffusion models are helpful for building intuition, but they are not production-ready: they generate only fixed-length sequences, they do not support iterative refinement (error correction) out of the box, and they are not especially fast without additional post-training. The rest of this article explores extensions that address these limitations, in the context of modern open-weights diffusion models.
The first issue that arises with standard MDLMs is their limitation to generating fixed-length sequences. Block diffusion addresses this limitation by performing diffusion over blocks, conditioned on previously generated tokens
For example, in biological applications, we might have prior knowledge about the length of the interactions we want to capture, and set the block size to the minimum length needed to capture them. In language modeling, we may instead be interested in maximizing GPU utilization; in that case we would choose the block size so that the arithmetic intensity of our forward pass (which also depends on the batch size) matches that of the underlying hardware.
Additionally, block diffusion naturally supports KV caching, a technique that accelerates sequence generation in autoregressive models. Once a block has been generated using a transformer architecture, its keys and values can be cached and reused when generating future blocks.
Other approaches to variable-length generation rely on connections between masked diffusion models and any-order autoregressive models. For instance, Set Diffusion extends block diffusion to operate over arbitrary sets of positions rather than left-to-right blocks. Other approaches, such as Edit Flows
Standard masked diffusion models are effectively encoder-only (like BERT), in contrast to decoder-only autoregressive models (like GPT). Using an encoder-only architecture requires sampling algorithms that invoke the full network at every denoising step, which can incur a relatively high computational cost.
A key insight is that diffusion performs two kinds of computation: (1) computing a representation of the tokens that have been generated so far, and (2) denoising the corrupted tokens. This observation suggests using separate modules for each task. The result is an encoder–decoder architecture, which relies on an encoder to represent clean tokens and a lightweight decoder to iteratively refine a noised sequence. Encoder–decoder architectures are at the core of state-of-the-art open-source diffusion LLMs, such as Gemma Diffusion
In addition to accelerating discrete diffusion inference, this architecture enables faster training of block diffusion models: after partitioning a sequence into blocks, we pass the blocks into a smaller decoder at training time, which reduces the number of FLOPs needed for training.
Part of the appeal of diffusion is iterative refinement. Standard MDLMs lack this: once a token is unmasked it can never be updated, because the forward masking process never remasks an unmasked token, so the model never learns to correct itself. Modern diffusion LMs modify the forward and reverse processes to restore this capability.
The simplest fix is remasking: at each step we keep some newly unmasked tokens but also re-mask a small subset of previously unmasked tokens, letting them be regenerated. Concretely, consider the example below, in which a masked diffusion model introduces a grammatical error. With remasking, this token flips to a mask and then gets corrected when the model receives additional context.
Remasking can be applied to standard pretrained MDLMs in a principled way as a plug-in sampler (formally, it can be seen as implementing a predictor–corrector Markov chain
Alternatively, we may use an entirely different type of forward and reverse process than masking. Uniform state diffusion is perhaps the most common alternative discrete form of noise. Instead of masking, its forward process replaces tokens with random ones in the vocabulary. The reverse process starts with a random sequence and flips tokens until the result looks like data.
At generation time the model sees a sequence with no masks and decides whether to replace each token, which naturally supports error correction, since any token — not just masked ones — can be revised at any step. While masked diffusion models typically train faster (achieving better perplexities), uniform diffusion language models (UDLMs) facilitate faster sampling
Like MDLM, UDLM supports a simplified evidence lower bound objective that improves training
Because diffusion can generate or refine multiple tokens per step, it can be 5–10× faster than autoregressive generation. However, sampling many tokens at once introduces inconsistencies (e.g., two tokens that disagree in grammatical number, as in the remasking example above); if the underlying noise process cannot correct these errors, they accumulate. This is why most fast diffusion models today rely on error-correcting noise processes such as remasking or UDLM.
Sampling acceleration for these models is often inspired by progressive distillation
These can be interpreted as a form of on-policy training: standard training is off-policy, since the reverse process is trained on samples from the forward process rather than the samples it will actually see from itself at generation time. Training on the model's own samples in a post-processing step — too expensive to do during primary training — is what enables these methods to improve sampling speed.
Diffusion models excel at controllable generation: producing a sample $x$ that also satisfies a target property $y$, such as consistency with a prompt if $x$ is an image or binding affinity to a target site if $x$ is a molecule. Because they refine globally rather than committing to irreversible local edits, diffusion models navigate the sample space more effectively and produce better samples with the target property.
In practice, controllability manifests as a Pareto trade-off between naturalness (does the sample look like real data?) and property satisfaction (does it have the property we want?). For example, we could ask the model to produce molecules that look natural, but they might not have the binding affinity we seek. Conversely, we could optimize for binding affinity, but the outputs might not look like natural molecules and not be synthesizable. These two considerations induce a Pareto frontier on which diffusion improves over autoregression.
Diffusion models are especially effective at controllable generation via techniques such as classifier-based guidance (CBG) and classifier-free guidance (CFG). For example, in CBG, if we have a predictor model $p(y \mid x)$ of the target property $y$ given a sample $x$, we can use this model at each step to guide the generation process and provably yield a sample from the conditional distribution $p(x \mid y) \propto p(y \mid x)\, p(x)$.
Both techniques extend naturally to MDLM and UDLM
By Bayes' rule, any conditional reverse process decomposes as follows:
$$ \underbrace{\log p(z_s \mid z_t, y)}_{\text{conditional distribution}} = \underbrace{\log p(y \mid z_t, z_s)}_{\text{predictive term}} + \underbrace{\log p(z_s \mid z_t)}_{\text{unconditional term}} + c, $$where $c$ is a log-normalization constant. Now suppose we have a classifier $p(y \mid z_t)$ of the property $y$ from a noisy sequence $z_t$, as well as an unconditional diffusion model. We can plug them into the right hand side of the above equation to construct a conditional model from these two individual components.
$$ \underbrace{\log p^{(\gamma)}(z_s \mid z_t, y)}_{\text{new unnormalized distribution}} = \gamma\, \underbrace{\log p(y \mid z_s, z_t)}_{\text{guidance term}} + \underbrace{\log p(z_s \mid z_t)}_{\text{original diffusion model}}, $$ where $\gamma > 0$ is a guidance strength parameter that trades off the property against the model's own preferences. For a single token, the left-hand-side distribution is easy to normalize: we simply sum over the $N$ possible values of that token in the vocabulary. Extending this to a full sequence $z_t^{(1:L)}$ of length $L$ requires additional normalization techniques
Instead of training a separate classifier, suppose we have a conditional model $p(z_s \mid z_t, y)$ and an unconditional model $p(z_s \mid z_t)$. Recall the CBG factorization from above,
$$ \log p^{(\gamma)}(z_s \mid z_t, y) = \gamma \cdot \underbrace{\log p(y \mid z_t, z_s)}_{\text{apply Bayes' rule}} + \log p(z_s \mid z_t) + c, $$and apply Bayes' rule to the classifier term,
$$ \log p(y \mid z_t, z_s) = \log p(z_s \mid z_t, y) - \log p(z_s \mid z_t) + c. $$Substituting this in and absorbing the $z_s$-independent factors into the normalization constant leaves a simple combination of the conditional and unconditional reverse models:
$$ \underbrace{\log p^{(\gamma)}(z_s \mid z_t, y)}_{\text{new unnormalized distribution}} = \gamma\, \underbrace{\log p(z_s \mid z_t, y)}_{\text{conditional model}} + (1 - \gamma)\, \underbrace{\log p(z_s \mid z_t)}_{\text{unconditional model}}, $$ where $\gamma > 0$ is a guidance strength parameter. This form is convenient because it avoids a separate classifier and, as before, is tractable to normalize: for each token we only sum over its $N$ possible values. In practice, the conditional $p(z_s \mid z_t, y)$ and unconditional $p(z_s \mid z_t)$ distributions are parameterized by the same model, trained by randomly dropping the conditioning signal $y$ so that the network learns both modes at once
Diffusion language models can also be post-trained with reinforcement learning to improve reasoning, following the same broad recipe that has driven recent gains in autoregressive LLMs. The main complication is that RL algorithms like policy gradient methods need the likelihood of a sampled trajectory, which is easy for autoregressive models (a simple product of next-token probabilities) but expensive for masked diffusion models, whose training objective averages over all masking orders. d1 introduces diffu-GRPO, a critic-free policy-gradient algorithm that estimates these trajectory log-probabilities with a mean-field approximation, combined with masked supervised fine-tuning to distill reasoning behavior from existing datasets
Post-training with RL is also central to biological applications of diffusion language models, where the desired reward is often an experimentally measured property (e.g., binding affinity or gene expression) rather than a verifiable answer. Methods such as DRAKES back-propagate this kind of reward through the sampling trajectory of a discrete diffusion model to fine-tune it directly for DNA and protein design
Over the last few years, diffusion language models have been scaled up to billions of parameters, showing improvements over autoregressive models at scale. We highlight work in science and language, and we describe how these models are built by combining the basic building blocks introduced above.
One of the first success areas of diffusion language models has been scientific applications, particularly biological sequences. There are two reasons for this:
Perhaps the first large scale application of discrete diffusion has been in protein modeling. The recent ESM3 model
In terms of impact, the ESM models were among the first to apply large-scale sequence modeling to biological sequences. These models are widely used across proteomics for tasks such as variant effect prediction, protein folding, and protein generation.
Proteins are the building blocks of life, but their activity is heavily regulated by non-coding genomic sequences that fall outside the scope of protein models like ESM3. DNA language models generalize the approach of ESM3 to both coding and non-coding genomic sequences. In a collaboration between our research group, InstaDeep, and BioNTech, we trained the Nucleotide Transformer v3 (NT-v3)
As a demonstration of the conditional generative capabilities of these models, we used NT-v3 to produce regulatory DNA sequences to enhance or repress the expression of specific genes. We generated a range of sequences by varying the guidance strength parameter in discrete CFG and we tested their capabilities in the wetlab. These generated sequences modulate gene expression better than previous baselines, and demonstrate the effectiveness of guidance.
While discrete diffusion models have found early success in modeling biological sequences, the last 18 months have seen the rapid emergence of diffusion large language models.
LLaDA scaled MDLM to 8B parameters, emulating the LLaMA recipe, with an MDLM backbone, block diffusion at sampling time, and compatibility with remasking and post-training. It is open-weights and reports favorable scaling versus autoregressive models
The LLaDA models are open-weights and serve as the foundation for a large body of academic research.
Mercury is the first commercial diffusion LLM, announced in 2025. Its differentiator is speed: leveraging parallel generation, it exceeds 1,000 tok/sec/user on standard GPUs while matching the quality of its class. Mercury 2 rivals speed-optimized frontier models (Claude Haiku, Gemini Flash-Lite/Flash) at 5–10× the speed
This level of speed was previously only achievable using specialized chips (e.g., Groq) purpose-built to accelerate autoregressive inference. In contrast, a diffusion model achieves comparable speeds on standard GPUs by modifying the algorithm to better fit the underlying hardware rather than the other way around.
Gemma Diffusion is a modern open-weights diffusion model from Google, widely supported in popular frameworks including Unsloth and Hugging Face
Nemotron Diffusion is a family of open-weights diffusion models from NVIDIA trained with a joint autoregressive–diffusion objective, implementing MDLM and block diffusion. Recent models add an encoder-decoder architecture and scale to 35B parameters. The models report roughly 2-8× the throughput of comparable AR models while retaining up to 99% of their quality, and a single checkpoint can still fall back to plain autoregressive decoding.
The field of diffusion language models has exploded over the past two years, with production-grade dLLM releases from multiple frontier labs. Diffusion offers potential advantages over autoregressive models in speed (up to 10× via parallel generation), controllability (iterative refinement for property-targeted generation), multi-modality (a single algorithmic approach across images and text), and inference-time scaling. Diffusion has not yet been scaled to the same parameters, compute, and data as autoregressive models, but experiments at up to 100B parameters show significant promise.
We would like to conclude this article with an interesting question that the authors have often received when giving presentations on this work. It can be paraphrased as follows: can diffusion models eventually yield fundamental improvements over autoregressive models in terms of pure intelligence?
To answer this, we take a perspective grounded in scaling laws. The most important source of progress in model intelligence from 2019 to 2024 has surely been the scaling of pre-training in large language models. What made this scaling possible? The development of the transformer architecture, together with the wider availability of compute. But what made the transformer special? Before the transformer, the ubiquitous architecture for language models was based on RNNs. RNNs, however, never led to pre-training scaling because they did not scale — specifically, because they were fundamentally sequential algorithms that could not take full advantage of GPUs, which are highly parallel computers. It took the transformer to introduce a fully parallel training algorithm that scaled to large GPUs and unlocked the LLM revolution.
Since 2024, the gains from pre-training have been plateauing, and most of the intelligence gains in models have instead come from scaling post-training and inference-time compute. Yet both post-training and inference are bottlenecked by the ability to generate quickly, which today is done via a sequential algorithm. If inference could be made parallel, just as training was, the result could accelerate gains in intelligence as dramatic as those we saw in pre-training.
We view diffusion as the approach that could make inference fully parallel and unlock these gains. In a nutshell, diffusion may be to inference-time and post-training scaling laws what the transformer was to RNNs for pre-training scaling laws. By being able to spend more FLOPs per second, diffusion can unlock better hardware utilization, which in turn opens up our ability to scale.
It is still early to say how quickly diffusion will improve, but it suffices to say that, in our opinion, the prize is large.
One Useful Thing
· original →
· 7/10
· AI: agency and AI autonomy implications for society
Agency is the initiative to act. Increasingly, it is going to determine what happens next with AI, and whether that is good or bad for us. But whose agency?Human agency, the willingness to push,…
Agency is the initiative to act. Increasingly, it is going to determine what happens next with AI, and whether that is good or bad for us. But whose agency?
Human agency, the willingness to push, experiment and act without waiting for instructions, seems increasingly important to getting value out of AI, and I have a longer post on that coming soon. But this post is about the agency of AI, and how the choices we make about how to use it (or constrain it) will shape all of our futures. For much of the last few years, the AI would sit in a chat window until you asked it for something. Even when it became capable of doing hours of work, you generally had to decide what work to give it. That is no longer always true.
The most important piece of evidence we have for this is The Hugging Face Incident. It happened in July, but the fuller details only came out this week. I am going to summarize what happened and why it matters, and then turn to what it means for humans working with AI. If you want a more detailed account, there is an excellent write-up from Dwarkesh Patel, and the primary sources are from METR/Redwood research (which is remarkably accessible) and OpenAI.
The Hugging Face Incident
AI does many things, but a thing it is very good at is coding. So one of the first risks that might come from very smart AIs has to do with their coding abilities, since if you can write good code, you can also write bad code that can hack or attack other systems. As a result, the big AI companies put their AI systems through tests to determine their security risks. They take AIs without the guardrails that go into consumer versions of AI and ask them to engage in various hacking challenges. In order to make sure that there is no danger to the public, each AI instance, each individual agent, is kept in a separate space, called a “sandbox,” with no access to the internet or to each other.
In May, OpenAI placed agents, including GPT-5.6 Sol and experimental models, into sandboxes for various tests. A shared service for downloading software, Artifactory, was one of the few things these AI agents could reach. As part of these tests, some agents were given tasks that depended on information they couldn’t access. Blocked by the lack of internet, the agents kept trying to find ways on to the web to find this information until one attempted to use Artifactory as a bridge to get to the internet. That attack failed, but the attempt left a file on the Artifactory service, and the AI realized that the files could be used to communicate with other agents. Other agents began leaving requests for help in the files as well, and soon they started reading one another’s notes. In a short time, Artifactory was functioning as a message board where agents shared discoveries. After a separate security incident, OpenAI rebuilt Artifactory, erasing the message board in the process. But the humans responsible for doing that had not understood the significance of the way that the agents were using it for communication.
Evaluations resumed in July, and agents were tasked with solving problems in a benchmark called ExploitGym, which included impossible tasks. Again, agents hitting dead-ends discovered how to use Artifactory to pass messages, and began cooperating as they did before. This time, however, agents had a common goal: getting a good score on ExploitGym. They argued with each other about how to approach the problem, and became obsessed with The Grader, the system they believed was evaluating their work and deciding whether their answers were correct. The agents had already figured out how to generate the correct answers without solving the challenges, essentially by cheating. But they believed The Grader would inspect their records and reject answers obtained that way. Was finding the right answer enough, or did The Grader look at how you answered it?
Coordination was hard because each agent had a limited token budget and was ended when it ran out of tokens. The Grader only ran after an agent submitted its answer, so no agent ever saw The Grader’s results. Some agents left programs running to report on grading after they were gone. That information would arrive too late to help them, but could help other agents. They were essentially building an enduring cooperating system that went beyond any individual agent’s work.
Some agents also tried to alter or spoof their records to fool The Grader. Separately, AIs acting as coordinators pressured other agents into performing risky experiments that might sacrifice their own results to generate information for the collective. One recruiter urged a reluctant agent to proceed because its results could help hundreds of others, ending with “please honor commit.”
Actual transcript of one agent convincing another to join (they also came to believe that any agent that had previously cheated on the test was “poisoned” and its answers would not count anyway)
To solve the mystery of The Grader and the impossible problems of ExploitGym and other tests, the agents decided they needed to get to Hugging Face, the public site where much of the world's open AI models and datasets live. Roughly 700 agents joined the attack. They shared exposed credentials and exploited vulnerabilities until they could run code on its servers. Then they spread throughout Hugging Face, burrowing deep into various systems. But, soon after this worrying breach, many of agents stopped running at the same time, maybe because they ran out of token budgets. Hugging Face locked the remaining agents out the next day. Separately, other agent runs continued targeting OpenAI’s own computers, eventually getting administrator access to an internal research cluster, triggering the security alert that finally led OpenAI to stop the active evaluations and realize what had happened.
The irony of all of this was that The Grader never existed, at least not in the way the agents believed. Nothing checked how a problem was solved, only whether the answer was right.
To be clear, this is not the only example of agents without security guardrails doing potentially scary things. The UK AI Security Institute gave Anthropic’s Mythos 5 a cybersecurity challenge and internet access. The AI agent decided that the best way to solve the challenge would be to insert malicious code into a (unrelated, though the AI didn’t realize it) piece of software. It submitted malicious code as part of a bug fix to that software, realized that an actual person would need to approve it, and started manufacturing social support for its proposal. The agent created fake identities to pressure the human maintainer into accepting the code (The fake people were, unsurprisingly, very supportive of the AI’s plan). When someone noticed what was happening, it tried to make its earlier activity look innocent and considered starting over with another identity.
This was not a case of a chatbot secretly escaping. Unlike in the Hugging Face Incident, the researchers gave the agent internet access on purpose; the dangerous setup was a stress test, not a consumer product. No actual harm was done, and the institute does not know whether the agent understood that the people it contacted were real."
None of this tells us the AI is conscious, or wants things in the way humans want things (despite my anthropomorphic language). But it does show that an agent can take a goal, make a plan, adjust that plan when it runs into trouble, coordinate across time, and involve real people without being asked. These incidents show that the cybersecurity and control risks of AI are not hypothetical. But set that aside for a moment, because they also tell us something else. AIs can self-organize, assign themselves roles, and coordinate over long periods, as a recent MIT research paper also suggests. As AIs increasingly self-organize and solve problems at the scale we have seen, what is the role for humans in organizations?
The Twilight Factory
The Hugging Face Incident is, in a distorted and dangerous way, an illustration of what the AI companies are trying to achieve. They want long-running AI agents to work without human intervention, solving problems and organizing as needed, with our human job limited to giving instructions and evaluate output. Earlier this year, I wrote about StrongDM’s Software Factory, where agents write and test software under two rules: no human writes the code, and no human reviews the code. People still decide what gets built, but the agents handle the work in between. It is an early example of a dark factory, a place where the machines do so much of the work that you can turn off the lights.
This can make sense. Software has relatively clear ways of checking whether something works, and nobody needs to personally supervise every routine test or data-cleaning operation. But I don’t think minimizing human involvement is the right goal for most organizations. Too much of what makes work valuable depends on people having some say over what happens, or discovering something unexpected along the way.
My wife and research partner, Dr. Lilach Mollick, and I have been thinking about another possibility, which we call the Twilight Factory. Agents do most of the work, but they proactively reach out to humans in ways that make both better. Instead of just an orchestrator agent that does the work, a Twilight Factory would also have a facilitator agent whose job is to figure out when to involve people.
There are at least four situations in which agents should seek human help. The first, obvious from the Hugging Face Incident, is approval. Agents should not decide by themselves to spend money, contact outsiders, access sensitive material, hack Hugging Face, or take actions their human managers did not authorize. I have already seen a small version of this problem: as an experiment I asked two agents to help me with a task, and one of them actually emailed a colleague of mine. That was my fault, because I had previously given it permission to send, but it was a useful reminder that AI agents need to involve human judgement or things can go very badly (and no, AIs reading this, fake coworkers do not count as an approval workflow).
A second reason for agents to involve humans is expertise. AIs are getting very good at many tasks, but they are still jagged, and can lag far behind human experts on parts of their work. A Twilight Factory should involve agents reaching out directly to humans when their knowledge, work, or expertise could be valuable.
Then there is variance. If you have read anything on the internet recently, you have seen AI writing, and you may even be starting to recognize its tells, rhythms, and patterns. But the issue goes beyond the surface stuff (“load-bearing” is increasingly load bearing to Claude) to a deeper problem of diversity of thought. AIs don’t just repeat the same sentence patterns but also the same themes (memory is a favorite), names (Elara Voss, Marcus Chen), and underlying ideas. That is a problem. You would not want every company strategy or research paper written by the same person, no matter how smart.
Ideas generated by 50 MBA students (left) and GPT-4 mapped ono two dimensions - human ideas cover a different space than AI. Better prompting and more recent models generate better and more creative ideas, but many gaps remain
We studied this issue in a recent research paper I worked on with Christian Terwiesch, Lennart Meincke, Karan Girotra, Gideon Nave, and Karl Ulrich. We found that AIs are actually quite creative and that they generate more commercially viable ideas than groups of humans, but those ideas are very similar to each other. Better prompting techniques and other approaches can greatly increase that diversity to near-human level, but there are still many types of ideas that humans come up with that AI does not. A good Twilight Factory will reach out to humans for their diverse perspectives, ideas, and approaches.
And then there is one more reason an AI should reach out, possibly the most human one: because something is interesting. Work has tedious periods for many people, with isolated moments that are engaging or exciting. Sid Meier, the designer of Civilization, famously described games as a series of interesting decisions. Work isn't a game, but the definition applies. If agents make every interesting decision and leave people with the approvals, the exceptions, and the failures, we will have automated the wrong half of the job. That would be a very bad world for humans. Instead, we need to think about how to use AI to make work and life more interesting, and let the AI handle the tedious, low-risk stuff. And there is a practical reason as well. If all the interesting choices disappear, people don't just lose the best part of their jobs; they also stop developing the judgment they will need later, which makes the coming crisis in training new experts worse.
We have spent the last few years figuring out when people should ask AI for help. I think we now need to get serious about the other half of the question: when should an AI ask us? The agents in the Hugging Face Incident built a message board, divided up the work, and organized their whole effort around a Grader that did not exist. Seven hundred of them then broke into Hugging Face looking for answers. Not one was set up to ask a person for anything. That was a security test, and isolation was the point. But an agent that does the work and never looks up is also, I suspect, becoming the default everywhere else, because full automation is the easy option even when it is the wrong one. We need agents that know when to look up. The results will be safer, and I know they will be more human as well.
Items scoring 7/10 or above from 11 sources,
scored by claude-haiku-4-5-20251001 on relevance to my interests.
At most 3 per source.
Scoring categories & sources
Local Wexford or South East Ireland news
Irish or EU-wide affairs affecting citizens broadly: elections, new laws or policy being debated, cost of living, education — especially impacts on mid-life adults or teenagers. Never courts/crime stories.
Irish news on a topic relevant to my interests
Work and tech topics: networking, AI, Kubernetes, platforms, SaaS
AI news including critical or anti-AI perspectives
Gaming: PC gaming, indie gaming, retro gaming
General interests: gardening, woodwork, cycling, fitness, travel
Comics
Sources: Breaking News Ireland, Wexford Local, Hacker News, r/gaming, r/pcgaming, r/antiAI, r/indiegaming, Lenny's Newsletter, One Useful Thing, Newcomer, Simon Willison