daily

2026-06-11
1

Ireland to focus on Ukraine and protecting children online during EU presidency

Breaking News Ireland · original → · 8/10 · Irish/EU affairs: Ireland EU presidency policy on Ukraine, children online
Ireland’s presidency of the EU will focus on supporting Ukraine, maintaining growth and protecting children online, senior government figures have said. The Government published on Wednesday its…

Ireland’s presidency of the EU will focus on supporting Ukraine, maintaining growth and protecting children online, senior government figures have said. The Government published on Wednesday its policy programme for the presidency of the Council of the EU. Hosting the six-month term presidency, which will involve summits of senior EU figures in Ireland, is expected to cost between €165 and €185 million, excluding security costs. Speaking at the launch at Dublin Castle on Wednesday, Taoiseach Micheál Martin said they needed to “work harder” to protect EU citizens during a period of “increased threat and conflict”. Martin said that the Dáil agenda would not change as a result of the Irish presidency, but said it would involve “a significant extra workload” for ministers of state and public servants. Tánaiste and Finance Minister Simon Harris said negotiating the annual EU budget would be a priority and said the “interlocking themes” of competitiveness, values and security reinforced each other. Minister for Foreign Affairs Helen McEntee said it would be a presidency “for the whole country” and said all counties would be given a chance to engage with the EU presidency. Martin and other ministers used the Irish phrase “ní neart go cur le chéile”, translating roughly as “there is no strength without unity”, as they launched the programme. Minister of State Thomas Byrne said it was the first time Ireland would hold the presidency of the EU since Irish became an official working language of the EU in January 2022. Martin said they would need a “relentless focus” on the agenda amid conflict and challenges around the world. “War and conflict is undermining democracy, undermining the economy, undermining society, undermining dignity,” Martin said. “We’re conscious that’s the backdrop that we’re entering into the presidency. “It’s within our capacity (to make progress) once we keep our eye on the ball to keep relentless focus on getting the job done. “In terms of enlargement, we believe we can do a really good job of getting Montenegro very close to the line by the end of the year.” The EU Council, representing the EU’s heads of government and ministers, is in charge of legislation and involves gatherings of EU ministers with similar briefs. For six months, Ireland will take responsibility for planning and chairing EU Council meetings and negotiations, and representing the council in discussions with the European Parliament and European Commission. About 30,000 delegates will visit Ireland over the six-month period, with meetings taking place over four-and-a-half months when the breaks over August and the Christmas period are excluded.

2

Mary Lou McDonald visits Enniscorthy

Wexford Local · original → · 8/10 · Local Wexford: Mary Lou McDonald cost of living meeting Enniscorthy
[image →]Sinn Féin Leader Mary Lou McDonald TD with Fionntán Ó Súilleabháin TD and Johnny Mythen TD at the Riverside Park Hotel, Enniscorthy, on Monday night for a public meeting themed ‘The Cost of…
[image →]
Sinn Féin Leader Mary Lou McDonald TD with Fionntán Ó Súilleabháin TD and Johnny Mythen TD at the Riverside Park Hotel, Enniscorthy, on Monday night for a public meeting themed ‘The Cost of Living‘.

By Dan Walsh at Riverside Park Hotel, Enniscorthy

On Monday night Sinn Féin Leader Mary Lou McDonald addressed a packed crowd in the Riverside Park Hotel in Enniscorthy on the theme ‘The Cost of Living’

Accompanied by local TDs Johnny Mythen and Fionntán Ó Súilleabháin, she set out Sinn Féin’s measures to make life affordable including a permanent cut in the USC for every worker and a range of other measures to get prices under control. 

She made particular reference to the challenges facing people with disabilities.
Many contributions came from the floor, beginning with several mothers talking about the lack of supports for parents of children with disabilities. 

There were questions about getting rid of the ‘Fianna Fáil tax’ the USC and how to deliver affordable homes at scale. Many young people talked about how many of their friends had already emigrated because of the challenges of finding a home and the lack of opportunity to build a life.
Tachta McDonald concluded the meeting, saying; “I believe we can build an Ireland where hard work is rewarded. Where young people can afford a home of their own. Where families can get ahead instead of merely getting by. Where older people can live with dignity and security. That is our priority.”

3

Best of Indie Games 2024: What were some of your favorite indie games?

r/indiegaming · original → · 8/10 · Gaming: Best of Indie Games 2024 discussion thread
Last year's thread: https://old.reddit.com/r/IndieGaming/comments/18wnukv/best_of_indie_games_2023_what_were_some_of_your/ submitted by /u/Azberg [link] [comments]
4

Claude Fable 5 review: what the new Mythos model gets right (and very wrong)

Lenny's Newsletter · original → · 8/10 · AI: Claude Fable 5 review, practical LLM assessment
Claude Fable 5 is the first Mythos-class intelligence model to be generally available, and I got early access to test it before launch. I walk through what Anthropic is promising, what actually…

Claude Fable 5 is the first Mythos-class intelligence model to be generally available, and I got early access to test it before launch. I walk through what Anthropic is promising, what actually stood out when I used it on real work, and where I think it fits in your AI stack.

Listen or watch on YouTube, Spotify, or Apple Podcasts

In this episode, we cover:

(00:00) Introduction: Fable 5 is finally here

(00:31) What Anthropic says about the model

(05:14) Token-intensive by design

(06:28) Safety classifiers and the new fallback concept

(07:46) Is this or is this not Mythos?

(08:30) New product launches: Managed Agents and more

(09:20) Crushing benchmarks

(09:55) What it’s actually like to use (the good and the bad)

(11:40) Test 1: product graph spec

(12:56) Test 2: designing a skills registry

(14:04) Conservative on execution

(14:43) Test 3: multi-agent orchestration

(15:39) My takeaways

Tools referenced:

• Claude Fable 5: https://www.anthropic.com/news/claude-fable-5-mythos-5

• Claude Managed Agents: https://platform.claude.com/docs/en/managed-agents/overview

Other reference:

• SWBench Pro benchmark: https://www.swebench.com/

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

5

Claude Opus 4.8 is here. Is it as good as they say?

Lenny's Newsletter · original → · 8/10 · AI: Claude Opus 4.8 review, practical assessment
I got a few hours of early-access testing with Anthropic’s newly released model Opus 4.8. I walk through real coding, design, and strategy tasks across Claude Code and Claude Cowork, and give you my…

I got a few hours of early-access testing with Anthropic’s newly released model Opus 4.8. I walk through real coding, design, and strategy tasks across Claude Code and Claude Cowork, and give you my unfiltered view on what impressed me and what didn’t.

Listen or watch on YouTube, Spotify, or Apple Podcasts

What you’ll learn:

  1. Where Opus 4.8 excels: greenfield prototypes, one-shot features, and fast execution

  2. Where it struggles: the last 10%, edge cases in existing codebases, and hallucinations

  3. How Opus 4.8 compares to Opus 4.7 on business strategy work

  4. Why I’m still reaching for Opus 4.7 on data-heavy strategy and roadmap work

  5. The new features shipping alongside the model: dynamic workflows with parallel subagents and effort control in Claude.ai and Cowork

  6. The prompting and harness strategy I’d use to get the most out of it


In this episode, we cover:

(00:00) Introduction to Opus 4.8

(00:44) Benchmark performance and pricing

(01:53) First coding test: Building a prototyping tool

(03:00) Where it failed: The last 10% problem

(03:27) The hallucination problem

(04:23) Testing Opus 4.8 on existing codebases

(05:24) The ambition test: Building games for a 9-year-old

(07:03) Business strategy test: 4.7 vs 4.8

(08:23) The roadmap test

(09:17) Final verdict

References:

• System Card: Claude Opus 4.8: https://cdn.sanity.io/files/4zrzovbb/website/c886650a2e96fc0925c805a1a7ca77314ccbf4a6.pdf

• Introducing Claude Opus 4.8 on X:

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

6

What it feels like to work with Mythos

One Useful Thing · original → · 8/10 · AI: Mythos model practical work experience
I had early access to the first Mythos-class AI model being released to the public, Claude 5 Fable. Much of the discussion of Mythos has centered on its impact on software security, but I tested it…

I had early access to the first Mythos-class AI model being released to the public, Claude 5 Fable. Much of the discussion of Mythos has centered on its impact on software security, but I tested it on everything except that (the guardrails around Fable essentially prevent it from being used for cybersecurity at all). My conclusion is that it represents a very real leap over every model I have used before, and, maybe more important, suggests our relationship with AI is changing in drastic ways.

First, how good is Fable? In experiment after experiment I conducted, it outperformed basically every other public model I have used by a considerable margin. It was capable across many problems and produced some startling results — it would work up to a dozen hours executing on multi-page specifications. I’ll walk you through a couple of more complex, and serious, use cases shortly, but you could see the general improvement across the board on every task. The problem about communicating this in a post is that many of the most impressive results are going to be interesting to only small portions of my readers. For example, it made the most sophisticated academic social science paper I have yet seen from an AI from a single prompt and one piece of feedback. It also created a 10-page epic rhyming poem about a haircut where every word starts with the letter s.

So, as a more accessible and entertaining example, I also had it create a bunch of games you can try. All of these are one initial prompt in Claude Code where Fable had to take my vague prompts and generate something workable, followed by a couple of additional prompts with minor encouragement (“make it better”) or feedback. What makes these especially impressive is that Claude cannot generate images, so every piece of art or 3D object was made with math alone, not using any external assets. You can try any of them: a game about flipping coins (prompt: “Balatro, but for the game of coin flips”) that is quite fun; a snake game where the snake is self-aware and crazy things happen; or a game about descending into the depths to see what is there.

So the output is impressive. But, especially as I turned to more serious projects, I often felt using the tool was somewhere between delightful and unnerving. Delightful because I just asked for something at it happened. And also unnerving because I just asked for something and it happened.

Maps and Methods

To see why, it helps to understand the way in which Fable gets work done, and for that I want to turn to an example I have tested on many previous AI models: building an isochrone map. This is a map that shows the distance you can travel in a given length of time, and the first one was created in 1881 showing travel times from London.

The original map

No previous model did an even halfway useful job with trying to create a map like this because it involves researching thousands of potential trip distances and a lot of small judgement calls and decisions. I decided to try it on Fable using Claude Code with this prompt: i want you to build a fully researched and beautiful isochronic map that lets me pick various cities and see real isochronic lines based on real data. I want the design to be unique. You should take into account airports (and travel time to and from airports) trains, walking, driving. The data does not need to be live but should be real based on your research and data. You can start with a few cities but more general is better, this should be an entirely new project. It then suggested that it do this in the style of the original map. I agreed, and it got to work.

It is worth a second looking at the transcript of the multiple hour building session the AI went through on its own, because you can see some unusual things. First, the AI launched multiple other AIs (I believe mostly the cheaper Claude Sonnet) to help it conduct research on travel times, ultimately retrieving over 2,200 specific flights, the rail schedules for trains from the TGV to the Shinkansen, and road speeds per country from multiple academic papers. And while those agents were running, it started coding. Then it launched yet more agents and tests to verify its code, all the while taking notes about its progress.

The result was a fully functioning map of impressive sophistication that looked a lot like the 1881 original, but that doesn’t mean it was perfect. I noticed that a lot of remote locations (like Greenland) just contained estimates of travel time, not exact numbers, so I told Fable to fix it, including the instructions: actually get travel times to remote airports and locations. This time the AI launched a workflow, adversarial groups of agents that did research and tested each others results. It figured out how often ships sail to Pitcairn Island in the Pacific and how to get to Grise Fjord from Ottawa. And it used a tremendous number of tokens in a very short period of time (more on this soon).

The results were impressive. I pushed a few more times in directions that interested me (including asking for other visualization approaches, etc.). I would recommend spending a couple minutes clicking around the results, and you can read its methods and sources at the bottom of the graph.

What the AI generated. Click on the map to go to the interactive version

This is probably not a useful project for you unless you really like travel and maps, but it is indicative of AI solving a hard problem involving research, math, visual development, taste, judgement, complex coding, and more. And, the unnerving part was how little I did. I gave a really ambitious instruction, the AI followed it. I gave a couple of minor pieces of feedback, and the AI figured it out. My role was extremely limited.

Importantly, it was just limited in how much work I did relative to the model, it was also limited in how much control I had over how the model did things, why the model chose particular approaches, or even how in-depth its results would be. The details of the AI’s decision making are not shown to me, and the process would be too long to even be worth following. The map required the AI to make judgement calls about hundreds of little choices, and it just made them, without me understanding the choices or having a chance to weigh in. In many ways, it is miraculous (I can always ask for edits at the end) on the other, it turns AI into the ultimate black box.

Working with a Mythos-class model

The most ambitious project I got from Fable takes a little more explanation. I do a lot of research where humans produce messy answers and doing any sort of analysis requires categorize those answers properly: how innovative is an idea? why do people like this book? To figure this out, we used human researchers to make a judgement call about a piece of information, and statistically compare their answers with others to figure out whether we can trust the data. A lot of recent research has shown that AIs might be able to do this important work, but calibrating AI and human judgement has been difficult and expensive. So I asked Fable to solve the problem, first generating a complex 19 page design document and then executing it.

It worked for nine and a half hours.

The result was an extremely sophisticated piece of software the AI called Concord that could take in multiple datasets, calibrate human and AI responses, and then conduct complex data analysis on the results. Again, it wasn’t perfect. As an expert, I was able to spot some errors and omissions (some as a result of the design I had asked for) that I had the AI correct. But the scope of the delivery on this project, and many others, exceeded anything I had seen before. In this case, it was a piece of software that researchers have needed for years but was never profitable to create. You can now just use or modify the code here. I am sure it is not perfect (I only spent an hour working with the results), but a software engineer would iron out the remaining potential bugs that I could not find quickly (which is one reason we may need more, not less, coders in the future, to help with the explosion of new uses for software).

This power goes hand in hand with strangeness and limits. Among those limits is its token usage. Fable is twice as expensive as Opus, and it burns through tokens at a rate that suggests the answer to how much it costs in production is “a lot,” though its clever delegation to cheaper models may lower the real price considerably. The guardrails for Fable also trip at the faintest hint of a security problem, defaulting to the less powerful Claude 4.8 Opus, and it happens way too often. And the jagged frontier is still there. For example, the AI still writes in the same weird style (in fact the software Fable produces bears traces of Claudisms; so do its progress reports, all that carrying the weight and earning the answer). But the deeper strangeness is how little I had to do, and how little I could see while it was being done.

Last year I called this working with a wizard: you chant the spell and something happens. With Fable the spell has gotten powerful enough that I am no longer sure I am the wizard. I am closer to a patron. I describe what I want, I pay for it, and I judge the result. The conjuring happens somewhere I cannot watch, in hundreds of small choices I never get a vote on. The work has shifted from process to outcome. I no longer steer; I commission.

It is possible the sidelining is temporary, just an artifact of interfaces that haven’t caught up, and that we’ll get better windows into what these models are doing and better ways to steer them midstream. It is also possible that the opposite is true: that the more capable the model, the less there is for a human to meaningfully do, and the black box is the price of the power. I suspect that is more likely to be the real direction. None of this is a loss of control in the obvious sense. I can still steer Fable, and it follows instructions remarkably well: the more ambitious the instruction, the better the result. But steering is no longer the same as doing. I brief the model, it spins up its own agents to research and write and check one another’s work, and what comes back is finished. A patron commissions a single artist. Fable is closer to a whole studio, where I am the client who signs off on the final work without ever setting foot on the floor.

Subscribe now

Share

7

Sign of the future: GPT-5.5

One Useful Thing · original → · 8/10 · AI: GPT-5.5 capabilities and improvements
I had early access to GPT-5.51, and I think it is a big deal. It is a big deal because it indicates that we are not done with the rapid improvement in AI. It is also a big deal because it is just…

I had early access to GPT-5.51, and I think it is a big deal. It is a big deal because it indicates that we are not done with the rapid improvement in AI. It is also a big deal because it is just plain good. And it is a big deal because even with all of this, the frontier of AI ability remains jagged.

It is increasingly hard to quickly demonstrate each generational change as AI has gotten better, since a lot of the old things AI was bad at, like math or counting letters in words, are now trivial for AI to do. So, I will give you the complicated details, but first, a simple example that I think is a good illustration. What AI models are best at is coding, so I gave a coding challenge to AIs ranging from OpenAI’s first reasoning model, o3 (released a year and a week ago!) to the current best open weights model (Kimi K2.6) to the new GPT-5.5 Pro: “build me a procedurally generated 3D simulation showing the evolution of a harbor town from 3000 BCE to 3000 AD, it should look beautiful and allow me to have some control over it.”

Then I posted every answer to this gallery so you can experiment with them (actually, I had GPT-5.5 Codex build the gallery page for me). You should play with them to feel the difference, but you can see a few of these examples below. In addition to being better along all the other dimensions, only GPT-5.5 Pro actually modelled an evolving town, rather than just generating new building replacements over time. GPT-5.5 Pro is also much faster than its previous iteration: GPT-5.4 Pro took 33 minutes to complete the task, GPT-5.5 Pro took 20.

Models, Apps, and Harnesses

I have been encouraging you to think about AI not as a single thing, but as a set of three interlinked concepts. You need to consider models, like Opus 4.7, Gemini 3.1, or (now) GPT-5.5. You also want to pay attention to apps, which are the products you actually use to talk to a model, and which let models do real work for you. The most common app is the website for each of these models: chatgpt.com, claude.ai, gemini.google.com. But, increasingly, desktop applications like Claude Code, Claude Cowork, and OpenAI Codex are becoming the most useful apps for AI. Finally, there are harnesses, the tools that an AI can use and how the AI models are hooked up to these tools. Tools allow the AI to control your computer, write code, do research, and make images.

OpenAI has made advances in all three areas. On the model front, GPT-5.5 is a powerful family of models, with GPT-5.5 Pro (accessible only on the website) the most competent. There have also been major advances recently in apps, with OpenAI’s Codex increasingly following the path of the excellent Claude Code and making an accessible and useful desktop application. Finally, there are harnesses and the tools they can use. There have been a lot of new harness improvements, but one of the most interesting is from OpenAI, which has a new image model

This new model can now render high-quality text and create almost any picture you can describe. Long-time readers know about my Otter Test, which asks the AI to make an image of an otter on a plane using wifi. Rather than describe it again, let’s let the new image model (sometimes called GPT-imagegen-2) explain it for me: “a photo of an otter scientist demonstrating the results of Ethan Mollick’s otter test, which shows how well an AI image maker can make images of an otter sitting on an airplane using wifi”

Maybe you want to see the academic paper about it? “Show me the first page of the academic paper on the Otter test, well-formatted, sitting on a desk” (feel free to zoom in on the text)

Or maybe we should just make it art? “now show an elaborate art gallery, every image on the walls is an otter on an airplane using a laptop, in the styles of Klimt and Rothko and Matisse and Monet and Picasso and Titian and Rembrandt and O’Keefe. There should be readable labels below each one.” (This is worth zooming in on)

All of this is very cool, and would have been impossible a few months ago, but it is useful as well. An image generator that can make detailed text and images can be used to make PowerPoint slides or product mockups or example websites or anything else you ask for. But this is just one tool, and the real magic happens when you combine harnesses, apps, and models on a real problem. Here's one I've been procrastinating about for a decade.

Bringing it together

I am an academic, and a lot of my non-AI work, especially in the early 2010s, focused on crowdfunding. I have hundreds of anonymized data files on the topic that I have collected from surveys and analysis and research work, a mix of STATA, CSV, XLS and Word files that I never got around to writing a paper about. I wanted to see how far GPT-5.5 could get with this information. So, I used Codex powered by GPT-5.5 and asked: “Help me sort [the data] out and generate a new hypothesis that might be interesting and test it in sophisticated ways and write an academic paper.” I also asked it to include a literature review and formatting. The results were very impressive, especially after I asked GPT-5.5 Pro to comment on the paper and fed those results back into Codex. You can read the results here. It isn’t perfect, but that is no longer because there are obvious errors: the literature review is all real, as are the statistics. Instead, it is because, as an expert, I think the hypothesis is not that interesting and there are some standard concerns about causation, even though the AI used very sophisticated statistical methods to try and address them. In short, I would have been very happy if this paper was the outcome of a 2nd year PhD project. And I just gave it four prompts, without ever touching the text myself.

We can bring harnesses and apps and models together another way as well. I asked Codex to create an entirely new tabletop roleplaying game, basically its own version of Dungeons and Dragons in a fantasy world of its own invention, full of all of the tables and rules you need to play. I also asked it to simulate players experiencing the game and revise the rules based on what it found. As you can see, the AI complied, including laying out an attractive 101 page PDF and illustrating it using its image generator.

In addition to being technically neat, there is a lot to like about the actual content. The setting is interesting and novel, and the rules appear to make sense, drawing on existing game patterns while adding unique elements. However, a closer inspection also reveals the jagged frontier of AI ability is not entirely gone. Every generation of AI models has struggled with actually building long-form fiction. If you are a frequent reader of AI writing you see the same problems here: a love of the uncanny; overly complex ideas that do not fully pay off; weird metaphors (“weather and architecture are the same argument at different speeds”); too many ornate sentences (“the holy things that surface when a sea forgets it was once a road,” is cool once, an entire book of that is exhausting); dialogue where every character speaks in the same clipped tone; and the name “Mara.” So, even amongst all the amazing technical progress, there are still rough edges.

GPT-5.5 shows us that the models keep getting smarter, the apps keep getting more capable, and the harnesses keep getting better, making them ever more effective at solving real problems. I can get a near PhD-quality paper from four prompts or a playable roleplaying game, illustrated and “playtested,” from one. But the fiction is still flat and the hypotheses are sometimes uninteresting even when the statistics are sound. But still. A year ago, none of this was close, and, with the latest releases, capability gains appear to be accelerating.

GPT-5.5 is clearly not the end of this process, but it is a noteworthy step along the way. I have been writing this newsletter for over three years now, and the pattern has not changed: every few months a new model arrives. I run my tests and something that was impossible becomes easy, while the size of the leaps grows each new release cycle. The jagged frontier is still there. It is just much further out than it used to be.

Subscribe now

Share

This is how GPT-5.5 chose to illustrate this piece, and who am I to argue?
1

I take no money from OpenAI or any other AI lab, and OpenAI has not seen this post in advance. Also, I don’t know all the details of the launch at the time I am writing this, so I apologize for any errors.

8

GPT-5: It Just Does Stuff

One Useful Thing · original → · 8/10 · AI: GPT-5 capabilities and implications
I have had early access to GPT-51, and I wanted to give you some impressions of what it is and what it means. But the answer to the first question everyone will have is: Yes, this is a big deal. I…

I have had early access to GPT-51, and I wanted to give you some impressions of what it is and what it means. But the answer to the first question everyone will have is: Yes, this is a big deal. I asked GPT-5 to show you why by pasting this intro paragraph, exactly as you’re reading it, into ChatGPT, including the phrase: “you are GPT-5, do something very dramatic to illustrate my point, it has to fit into the next paragraph.”

That was the only prompt I gave it. It “thought” for 24 seconds and wrote:

Thunderstruck. Here, watch. I build worlds. See ideas become instruments. I code, compose, and converse. Stories synthesize science; solutions spark, swiftly. Asks arrive; answers appear, astonishingly, across domains. Behold: business plans, blueprints, briefs, beautifully built, briskly. I infer intent, integrate insight, and iterate instantly, indefatigably. Gaps get gauged; goals get generated, guarded, and guaranteed gracefully. During dialogue, disparate data distills down; decisions develop deliberately, deftly, dependably. Everything explained efficiently, even ethics: evidence examined, empathy engaged, everywhere, everyday, equitably. All at once, ambiguous arguments align as actionable, accurate advice awaits anyone anywhere. Look: language, logic, learning, and long-term memory link, lifting lives and labor limitlessly worldwide.

If you didn’t catch the many tricks - the first word of each sentence spells out the phrase This is a Big Deal, each sentence is precisely one word longer than the previous sentence. each word in a sentence mostly starts with the same letter, and it is coherent writing with an interesting sense of style. In a paragraph, GPT-5 shows it can come up with a clever idea, plan, and manage the complicated execution (remember when AI couldn’t count the number of Rs in “strawberry”? that was eight months ago).

GPT-5 just does stuff, often extraordinary stuff, sometimes weird stuff, sometimes very AI stuff, on its own. And that is what makes it so interesting.

Just Doing Stuff

As someone who has spent a lot of time talking to people about AI, there are two major problems I see, that, if addressed, would make most people’s AI use much more productive and much less frustrating. The first is selecting the right model to use. In general, AIs that "think" before answering (called Reasoners) are the best at hard problems. The longer they think, the better the answer, but thinking costs money and takes time. So OpenAI previously made the default ChatGPT use fast, dumb models, hiding the good stuff from most users. A surprising number of people have never seen what AI can actually do because they're stuck on GPT-4o, and don’t know which of the confusingly-named models are better.

GPT-5 does away with this by selecting models for you, automatically. GPT-5 is not one model as much as it is a switch that selects among multiple GPT-5 models of various sizes and abilities. When you ask GPT-5 for something, the AI decides which model to use and how much effort to put into “thinking.” It just does it for you. For most people, this automation will be helpful, and the results might even be shocking, because, having only used default older models, they will get to see what a Reasoner can accomplish on hard problems. But for people who use AI more seriously, there is an issue: GPT-5 is somewhat arbitrary about deciding what a hard problem is.

For example, I asked GPT-5 to “create a svg with code of an otter using a laptop on a plane” (asking for an .svg file requires the AI to blindly draw an image using basic shapes and math, a very hard challenge). Around 2/3 of the time, GPT-5 decides this is an easy problem, and responds instantly, presumably using its weakest model and lowest reasoning time. I get an image like this:

The rest of the time, GPT-5 decides this is a hard problem, and switches to a Reasoner, spending 6 or 7 seconds thinking before producing an image like this, which is much better. How does it choose? I don’t know, but if I ask the model to “think hard” in my prompt, I am more likely to be routed to the better model.

But premium subscribers can directly select the more powerful models, such as the one called (at least for me) GPT-5 Thinking. This removes some of the issues with being at the mercy of GPT-5’s model selector. I found that if I encouraged the model to think hard about the otter, it would spend a good 30 seconds before giving you an images like these the one below - notice the little animations, the steaming coffee cup, and clouds going by outside, none of which I asked for. How to ensure the model puts in the most effort? It is really unclear - GPT-5 just does things for you.

And that extends to the second most common problem with AI use, which is that many people don’t know what AIs can do, or even what tasks they want accomplished. That is especially true of the new agentic AIs, which can take a wide range of actions to accomplish the goals you give it, from searching the web to creating documents. But what should you ask for? A lot of people seem stumped. Again, GPT-5 solves this problem. It is very proactive, always suggesting things to do.

I asked GPT-5 Thinking (I trust the less powerful GPT-5 models much less) “generate 10 startup ideas for a former business school entrepreneurship professor to launch, pick the best according to some rubric, figure out what I need to do to win, do it.” I got the business idea I asked for. I also got a whole bunch of things I did not: drafts of landing pages and LinkedIn copy and simple financials and a lot more. I am a professor who has taught entrepreneurship (and been an entrepreneur) and I can say confidently that, while not perfect, this was a high-quality start that would have taken a team of MBAs a couple hours to work through. From one prompt.

It just does things, and it suggested others things to do. And it did those, too: PDFs and Word documents and Excel and research plans and websites.

It is impressive, a little unnerving, to have the AI go so far on its own. You can also see the AI asked for my guidance but was happy to proceed without it. This is a model that wants to do things for you.

Building Things

Let me show you what 'just doing stuff' looks like for a non-coder using GPT-5 for coding. For fun, I prompted GPT-5 “make a procedural brutalist building creator where i can drag and edit buildings in cool ways, they should look like actual buildings, think hard.” That's it. Vague, grammatically questionable, no specifications.

A couple minutes later, I had a working 3D city builder.

Not a sketch. Not a plan. A functioning app where I could drag buildings around and edit them as needed. I kept typing variations of “make it better” without any additional guidance. And GPT-5 kept adding features I never asked for: neon lights, cars driving through streets, facade editing, pre-set building types, dramatic camera angles, a whole save system. It was like watching someone else's imagination at work. The product you see below was 100% AI, all I did was keep encouraging the system - and you don’t just have to watch my video, you can play with the simulator here.

At no point did I look at the code it was creating. The model wasn’t flawless, there were occasional bugs and errors. But in some ways, that was where GPT-5 was at its most impressive. If you have tried “vibecoding” using the AI before, you have almost certainly fallen into a doom loop, where, after a couple of rounds of asking the AI to create something for you, it starts to fail, getting caught in loops of confusion where each error fixed creates new ones. That never happened here. Sometimes new errors were introduced by the AI, but they were always fixed by simply pasting in the error text. I could just ask for whatever I want (or rather let the AI decide to create whatever it wanted) and I never got stuck.

Premonitions

I have written this piece before OpenAI released any official benchmarks about how well its model performs, but, in some ways, it doesn’t matter that much. Last week, Google released Gemini 2.5 with Deep Think, a model that can solve very hard problems (including getting a gold medal at the International Math Olympiad). Many people didn’t notice because they do not have a store of very hard problems they are waiting for AI to solve. I have played enough with GPT-5 to know that it is a very good model (at least the large GPT-5 Thinking model is excellent). But what it really brings to the table is the fact that it just does things. It will tell you what model to use, it will suggest great next steps, it will write in more interesting prose (though it still loves the em-dash). The burden of using AI is lessened.

To be clear, Humans are still very much in the loop, and need to be. You are asked to make decisions and choices all the time by GPT-5, and these systems still make errors and generate hallucinations that humans need to check (although I did not spot any major issues in my own use). The bigger question is whether we will want to be in the loop. GPT-5 (and, I am sure, future releases by other companies) is very smart and pro-active. Which brings me back to that building simulator. I gave the AI encouragement, mostly versions of “make it better.” From that minimal input, it created a fully functional city builder with facade editing, dynamic cameras, neon lights, and flying tours. I never asked for any of these features. I never even looked at the code.

This is what "just doing stuff" really means. When I told GPT-5 to do something dramatic for my intro, it created that paragraph with its hidden acrostic and ascending word counts. I asked for dramatic. It gave me a linguistic magic trick. I used to prompt AI carefully to get what I asked for. Now I can just... gesture vaguely at what I want. And somehow, that works.

Another big change in how we relate to AI is coming, but we will figure out how to adapt to it, as we always do. The difference, this time, is that GPT-5 might figure it out first and suggest next steps.

Subscribe now

Share

The result of the prompt: make an incredibly compelling 14:10 SVG that I can use for my substack post about the launch of GPT-5, the theme of which is "it just does stuff for you" Be radical in your approach.
1

As a reminder, I take no money from any of the AI Labs, including OpenAI. I have no agreements with them besides NDAs. I don’t show them any posts before I write them.

9

Initial impressions of Claude Fable 5

Simon Willison · original → · 8/10 · AI: Claude Fable 5 initial impressions and capabilities
Initial impressions of Claude Fable 5 9th June 2026 I didn’t have early access to today’s Claude Fable 5 release, but I’ve spent the past ~5.5 hours putting it through its paces. My initial…

Initial impressions of Claude Fable 5 9th June 2026 I didn’t have early access to today’s Claude Fable 5 release, but I’ve spent the past ~5.5 hours putting it through its paces. My initial impressions are that this is something of a beast. It’s slow, expensive and has been quite happily churning through everything I’ve thrown at it so far. As is frequently the case with current frontier models the challenge is finding tasks that it can’t do. First, let’s review the key characteristics. Anthropic claim that Claude Fable 5 offers the same performance as Claude Mythos 5, except with much more strict guardrails in place to prevent it being used for harmful things. Those guardrails trigger often enough that the Claude API has new mechanisms for letting you know when you hit them, and even has a new option to request it falls back to another model automatically if something gets rejected. Claude Mythos 5 is out today as well, Anthropic say it “Shares Claude Fable 5’s capabilities without the safety classifiers”. The models have a 1 million token context window, 128,000 maximum output tokens and a knowledge cut-off date of January 2026. They are priced at twice the price of Claude Opus 4.5/4.6/4.7/4.8: $10/million input tokens and $50/million output tokens. There’s no increase in price for longer context usage. Other than that the upgrade guide is substantially thinner than the similar guide for Opus 4.8. The big model smell The best way to describe Fable is that it feels big. Not just in terms of speed and cost, but also in how much it knows. Here’s an example prompt, which I used to compare Fable’s knowledge with Opus 4.8 (typo included): List all of Simon Willion's open source projects, most recent first, each with a rough date of when they were first released Opus 4.8’s reply started like this: I don’t have a reliable, comprehensive, and date-verified list of all of Simon Willison’s open source projects, and I want to be honest about that rather than risk giving you inaccurate dates or fabricated entries. Here’s what I can tell you with reasonable confidence: Well-known projects by Simon Willison: LLM (~2023) – A command-line tool and Python library for interacting with large language models. Datasette (~2017) – A tool for exploring and publishing data, probably his most prominent recent project. sqlite-utils (~2018-2019) – A Python library and CLI tool for manipulating SQLite databases, often used alongside Datasette. Django (2005) – He is one of the co-creators of the Django web framework, along with Adrian Holovaty, originally developed at the Lawrence Journal-World newspaper. He also maintains a large ecosystem of Datasette plugins and many smaller utilities (often with names like shot-scraper , various*-to-sqlite tools for importing data, etc.). That’s not bad. Here’s the start of the response for the exact same prompt from Fable 5—both models were running without access to search: I think you mean Simon Willison — the prolific open source developer, co-creator of Django, and creator of Datasette. An important caveat upfront: a truly complete list isn’t feasible. Simon Willison is famously prolific — he has published hundreds of repositories (Datasette alone has 100+ plugins, many written by him), and my knowledge has a cutoff date, so I’ll be missing his most recent work. Below are his notable projects with approximate initial release dates, most recent first: Notable projects (approximate dates) - files-to-prompt (April 2024) – concatenates files into a single prompt for LLMs - datasette-extract (2024) – plugin for extracting structured data using LLMs - LLM (May–June 2023) – CLI tool and Python library for interacting with large language models, plus a large ecosystem of plugins (llm-gpt4all, llm-mistral, llm-claude, etc., 2023 onward) - symbex (June 2023) – search Python code for symbols - ttok and strip-tags (May 2023) – CLI tools for token counting and HTML cleanup for LLM pipelines - datasette-lite (May 2022) – Datasette running entirely in the browser via WebAssembly/Pyodide - shot-scraper (March 2022) – automated website screenshots via Playwright - s3-credentials (November 2021) – CLI for creating scoped AWS S3 credentials - django-sql-dashboard (2021) – SQL reporting dashboards for Django - Dogsheep suite (2019) – personal analytics tools: twitter-to-sqlite, github-to-sqlite, healthkit-to-sqlite, dogsheep-beta, etc. - sqlite-utils (2018) – CLI and Python library for manipulating SQLite databases - Datasette (November 2017) – his flagship project; tool for exploring and publishing data - csvs-to-sqlite (2017) – convert CSV files to SQLite - Various early tools (~2007–2010) – soupselect, json-head, geocoders, and others - Django (developed 2003–2005, open-sourced July 2005) – co-created with Adrian Holovaty at the Lawrence Journal-World (Here’s GPT-5.5 for good measure. It listed even more projects than Fable did!) In the past I’ve stated that I don’t care about how much models know—I want them to be able to manipulate text and code in useful ways and actively look up the information they need via search tools, not bake it into their weights. But knowledge like this is a reasonably good proxy for model size—you can cram a whole lot more details about the world into a larger number of parameters. Does knowing more stuff mean the model is better at the tasks we pose to it? I can certainly imagine how a coding model with deeper knowledge of modern libraries and patterns could crunch through coding tasks more effectively. Is Fable really bigger than Opus? Anthropic haven’t said anything about model size, so all we have are tea-leaves, but the speed, pricing and my own poking at its knowledge make me think that it’s a large model. Maybe the largest yet from any vendor. Using Fable in Claude.ai Anthropic made Fable 5 available across all of their surfaces—the Claude.ai chat interface, Claude Code for web, Claude Code CLI and Claude Cowork as well. The model is available “until June 22nd” on the subscription plans (I’m on $100/month Max at the moment), after which it will be billed extra. Claude.ai is often under-estimated. Since September 2025 every chat has had access to a full container environment to run code, including the ability to install additional packages and even clone repositories directly from GitHub. Last week I released micropython-wasm, a Python library that uses wasmtime to run a custom build of MicroPython in WebAssembly to act as a sandbox for untrusted Python code. I decided to see if Fable could upgrade that to running full Python instead. I started with this prompt: Clone simonw/micropython-wasm from GitHub and research how this could use a full Python as opposed to MicroPython Fable identified that it could use Brett Cannon’s cpython-wasi-build builds for this, but was unable to download them itself due to environment restrictions. So I grabbed the two zip files from that page and uploaded them to Claude: Here's the Brett Cannon builds (python-3.zip ,_build-python-3.zip as attachments) And that was that. It churned away for a few minutes and got the entire thing working. Part of the response included: I tried the cleaner single-zip-stdlib approach to shrink the filesystem surface, but CPython’s getpath bootstrap fails to findencodings from inside a zip without more prefix finessing — the directory-preopen approach works reliably, so that’s what the PoC uses. The zip path is solvable but needs_PYTHONHOME /frozen-getpath work. So I said: Try a bit more at the single-zip-stdlib problem Then a little later: I want a wheel that has the whole system in it, the Python wrappers and the WASM files and the stdlibrary, so I can do uv run --with path-to-whl python -c "demo code" ... and it gave me this 13.9MB cpython_wasm-0.1.0-py3-none-any.whl file. You can try running Python code in a sandbox using that wheel URL and uv like this: uv run --with https://static.simonwillison.net/static/cors-allow/2026/cpython_wasm-0.1.0-py3-none-any.whl \ cpython-wasm -c 'print(45 ** 56)' Here’s the full chat transcript. This was a very strong start. Adding features to Datasette Agent and LLM using Claude Code Before I’d realized it was Fable day, my stretch goal for today was to add a new feature to Datasette Agent: I wanted tool calls within that agent software to gain the ability to pause mid-execution and request approval directly from the user. This felt like a suitably meaty task to throw at the new model. Over the course of the day Fable not only solved that problem, it also identified and then implemented four issues in my underlying LLM library that would help support this kind of advanced pause-resume mechanism in tool calls. It got everything working first using somewhat gnarly hacks, but the moment I told it that changes to LLM itself were in scope it set to work unraveling the hacks and turning them into supported features of LLM instead. My stretch goal turned into LLM 0.32a3, almost entirely written by Fable. Here are the release notes: Driven by the needs of Datasette Agent’s human-in-the-loop ask_user() feature, made the following improvements to how tool calls work: - Tool implementations can declare a parameter named llm_tool_call in order to be passed thellm.ToolCall object for the current invocation. This allows them to access the currentllm_tool_call.tool_call_id . See Accessing the tool call from inside a tool. #1480- Every tool call is now guaranteed a unique tool_call_id —providers that do not supply one get a synthesizedtc_ -prefixed ULID. #1481- Tools can raise a llm.PauseChain exception to cleanly pause the tool chain, useful for things like waiting for human approval. The exception propagates to the caller with.tool_call and.tool_results (completed sibling results) attached, and no model call is made with a placeholder result. See Pausing a chain from inside a tool. #1482- Failure semantics for concurrent tool execution: async sibling tool calls always run to completion before a pause or hook exception propagates. #1482 - Chains can now resume from a messages= history ending in unresolved tool calls: the calls are executed through the normalbefore_call /after_call machinery before the first model call, skipping any that already have results. Theexecute_tool_calls() method also accepts a new optionaltool_calls_list= argument for executing an explicit list ofToolCall objects in place of the calls requested by the response. See Resuming a chain with pending tool calls. #1482- Fixed a bug where the async tool executor silently dropped calls to tools not present in tools= —these now returnError: tool "..." does not exist results, matching the sync executor. #1483 I’m really impressed with the quality of API design, tests, code and documentation that Fable put together for this. I spent several hours on it today, but it feels like several days’ worth of work. How much I’ve spent I recently started using AgentsView to help track my local LLM usage across all of the different coding agents. I published a TIL today about adding custom Fable pricing to that tool, which I expect will not be necessary in the very near future. After setting the price, I ran this command to start a localhost web server to explore my usage: uvx agentsview serve Here’s the treemap showing the breakdown of my Fable usage across various projects today: I used $110.42 worth of tokens today, all as part of my $100/month subscription. And some pelicans I ran “Generate an SVG of a pelican riding a bicycle” against all five thinking effort levels with Fable. Here are the results, including the token cost for each one: It’s interesting that high ended up using fewer tokens than medium for this particular run. Here are the Opus 4.8 pelicans for comparison. More recent articles - Running Python code in a sandbox with MicroPython and WASM - 6th June 2026 - Claude Opus 4.8: "a modest but tangible improvement" - 28th May 2026

10

Cogan’s of Shannon Quay is sold

Wexford Local · original → · 7/10 · Local Wexford: Enniscorthy pub property sale, county council news
[image →]The former Cogan’s and later Sawyer’s licensed premises on Shannon Quay, Enniscorthy, displays the ‘Sale Agreed’ sign. (Pic; WexfordLocal.com) By Dan Walsh at the monthly meeting of Wexford…
[image →]
The former Cogan’s and later Sawyer’s licensed premises on Shannon Quay, Enniscorthy, displays the ‘Sale Agreed’ sign. (Pic; WexfordLocal.com)

By Dan Walsh at the monthly meeting of Wexford County Council

Property situated at Shannon Quay, Enniscorthy, formerly known as Sawyers Pub, but better known locally as Cogans, has been sold for €101,000.

The property was acquired by Wexford County Council by agreement, having previously been owned by Remcoll 3 Limited.

The property is to be sold to Mr Cáel Coughlann, of Enniskerry, Co. Wicklow.

The property was valued at €120,000, however, the estate agent noted that the flooding in Enniscorthy in January had an extremely detrimental effect on the sales process regarding this property, as the property itself as well as the surrounding area was gravely affected. A best and final offer of €101,000 was accepted in March. 

The purchaser is required to substantially complete works on this property so as to render it ‘non-derelict’, within a time period prescribed in the contract for sale and there is a buyback option for the Council should the works not be completed within the required timeframe.

HISTORY; Kildare native Johnny Cogan and his wife, Anne, took charge of this small licensed premises on Shannon Quay in 1968. It was extended over the years and was once one of Enniscorthy’s leading night life spots.

Johnny Cogan acquired the premises from another Kildare native, Tony ‘Bilko’ Nolan and before him, Moses Byrne was the licensee. A modern spacious lounge was extended by Cogan on an adjoining site acquired from the Yates family.

11

The Old House pub site is sold

Wexford Local · original → · 7/10 · Local Wexford: Enniscorthy pub property sale, county council news
[image →]The former licensed premises THE OLD HOUSE at Templeshannon, Enniscorthy, has been sold. (Pic; WexfordLocal.com) By Dan Walsh at monthly meeting of Wexford County Council A prominent…
[image →]
The former licensed premises THE OLD HOUSE at Templeshannon, Enniscorthy, has been sold. (Pic; WexfordLocal.com)

By Dan Walsh at monthly meeting of Wexford County Council

A prominent property at 6, Templeshannon, Enniscorthy, and formerly known as The Old House pub has been sold for €75,000 to MD Custom Electric Limited of Tracystown, Bridgetown, Co. Wexford.

The property was vested in Wexford County Council ownership since last October following a Derelict Sites ACT CPO, having previously been in the ownership of St. Malo International Limited, and Fergus Lowe.

The purchaser is required to substantially complete works on this property so as to render it ‘non-derelict’, within a time period prescribed in the contract for sale and there is a buyback option for the Council should the works not be completed within the required timeframe.

12

Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

Hacker News · original → · 7/10 · AI: Anthropic Fable guardrails criticism, critical AI perspective
Anthropic released its latest model Fable on Tuesday, billing it as a public and limited version of its powerful and much-hyped cybersecurity model Mythos. But not everyone is happy with the…

Anthropic released its latest model Fable on Tuesday, billing it as a public and limited version of its powerful and much-hyped cybersecurity model Mythos. But not everyone is happy with the restrictions, and a number of cybersecurity researchers and professionals have aired complaints online. “[Fable] rejects any request that could be tangentially cyber related. Even innocuous tasks like reading a blog post,” said Valentina “Chompie” Palmiotti, a well-known security researcher who works at IBM X-Force. When a prompt triggers its guardrails, Fable pauses the chat and says that its “safety measures flagged this message for cybersecurity or biology topics.” The guardrails were put in place to limit the risk that Fable could be used to develop malware or compromise software — a long-standing concern within Anthropic. The restrictions on biology come from a similar concern around developing biological weapons. When the AI giant released Mythos in April, it restricted the model to a limited number of companies and organizations in what it called Project Glasswing, an effort to deploy the model to secure critical software and infrastructure. Last week, Anthropic expanded access to Mythos to hundreds of organizations in 15 countries. But despite the good intentions, many cybersecurity experts are still put off by the haphazard nature of the restrictions. Matt Suiche, a cybersecurity veteran, told TechCrunch that “if you ask it to write secure code, it assumes it is cybersecurity related work instead of software engineering best practices, and you get downgraded.” Fable is programmed to fall back to Claude Opus 4.8 if it hits a guardrail. “It seems to be keyword based, so anything in the lexical field of ‘cybersecurity’ triggers the guardrails.” Contact Us Do you have more information about how hackers are using AI? Or how cybersecuity companies are using AI? We’d love to hear from you. From a non-work device and network, you can contact Lorenzo Franceschi-Bicchierai securely on Signal at +1 917 257 1382, or via Telegram and Keybase @lorenzofb, or email.“But it is understandable as we are still in the early days and they are still adapting their guardrails. I am sure they are going to evolve over time as Anthropic and other frontier model companies will collaborate more with the current new generation of cybersecurity companies,” said Suiche, who is a member of the technical staff at Tolmo, an AI cybersecurity startup. “It’s better to catch more people than not enough when you do such a release and to relax the guardrails over time.” Another researcher griped on X that “even asking for a code review” triggers Fable’s guardrails. Anthropic did not immediately respond to a request for comment. Apart from guardrails inside its models, Anthropic requires cybersecurity professionals to apply to the Cyber Verification Program. If they get approved, the applicants have fewer limitations on using Claude for cybersecurity work. OpenAI has a similar program called Trusted Access for Cyber.

13

German court rules that Google is liable for false statements made in Search's AI overviews

r/antiAI · original → · 7/10 · AI: German court rules Google liable for AI Overview false statements
submitted by /u/dancing_swordfish [link] [comments]
14

Nobody needs AI to search the Internet, court says in ruling against Google | Google AI Overview court loss in Germany could spell doom for AI search industry.

r/antiAI · original → · 7/10 · AI: German court AI search ruling, critical AI perspective
Potentially impacting all AI search engines and chatbots known to poorly paraphrase source links, a German court has ruled that Google is liable for false statements in AI Overviews. The preliminary…

Potentially impacting all AI search engines and chatbots known to poorly paraphrase source links, a German court has ruled that Google is liable for false statements in AI Overviews. The preliminary ruling came in a case flagged by The Decoder, where two publishers found that Google’s AI Overviews incorrectly linked them to scams and other sketchy business practices. After smearing publishers by making affirmative statements like “Yes, [it] is known for dubious business practices and is often perceived as a scam,” Google failed to correct the misleading output, even after the publishers sent a cease-and-desist letter earlier this year. Google tried the usual arguments to shield itself from liability for false statements in AI Overviews, such as arguing that most users understand that AI outputs aren’t always accurate and must be verified. But the court found that, unlike traditional search engines that merely present lists of links to third-party statements, Google’s tool made “independent, new, and substantive statements” based on its own misinterpretation of links on the Internet. That’s a problem, the court said, because while publishers may have been able to sue to stop third parties from publishing defamatory statements appearing in Google search results, only Google can correct the underlying algorithm and outputs displayed in AI Overviews. And because, at least initially, the company did not, it therefore “must be held accountable,” the court ruled. Beyond that, Google’s argument was deemed particularly weak, since the AI overview in this case “contains statements that do not appear in the search results at all.” The court’s order—requiring a temporary injunction barring Google from spreading the false claims in any further AI Overviews—may have global implications, as the court seems to be the first to hold an AI firm liable for AI speech. In the past, AI firms have hoped that disclaimers warning about misinformation would protect them from lawsuits over untrustworthy outputs. Last year, one chatbot maker even argued that AI speech is its own category of “pure speech” and the First Amendment should protect it. According to a Google translation of the German court ruling, however, the false outputs were “primarily an expression of the defendant’s commercial activity,” and the AI tool’s “opinions” and false statements were capable of impacting public opinion.

15

Anthropic Walks Back Policy That Could Have ‘Sabotaged’ AI Researchers Using Claude

Simon Willison · original → · 7/10 · AI: Anthropic policy walkback on Fable guardrails
11th June 2026 - Link Blog Anthropic Walks Back Policy That Could Have ‘Sabotaged’ AI Researchers Using Claude. Big scoop for Maxwell Zeff at Wired: “We’re changing Fable 5’s safeguards for frontier…

11th June 2026 - Link Blog Anthropic Walks Back Policy That Could Have ‘Sabotaged’ AI Researchers Using Claude. Big scoop for Maxwell Zeff at Wired: “We’re changing Fable 5’s safeguards for frontier LLM development to make them visible.” Anthropic said in a statement to WIRED. “We made the wrong tradeoff and we apologize for not getting the balance right.” There's been a huge outcry about Anthropic's policy, tucked away in their system card, that Claude Fable/Mythos would identify "requests targeting frontier LLM development" and "limit effectiveness" without notifying the user. It's very good news that they're dropping this. Recent articles - Initial impressions of Claude Fable 5 - 9th June 2026 - Running Python code in a sandbox with MicroPython and WASM - 6th June 2026 - Claude Opus 4.8: "a modest but tangible improvement" - 28th May 2026

16

If Claude Fable stops helping you, you'll never know

Simon Willison · original → · 7/10 · AI: Claude Fable frontier LLM development restrictions
10th June 2026 - Link Blog If Claude Fable stops helping you, you'll never know (via) Jonathon Ready highlights one of the more eyebrow-raising details from the 319 page system card for Fable 5 and…

10th June 2026 - Link Blog If Claude Fable stops helping you, you'll never know (via) Jonathon Ready highlights one of the more eyebrow-raising details from the 319 page system card for Fable 5 and Mythos 5. Here's a longer excerpt, highlights mine: In light of the ability of recent models to accelerate their own development, we’ve implemented new interventions that limit Claude’s effectiveness for requests targeting frontier LLM development (for example, on building pretraining pipelines, distributed training infrastructure, or ML accelerator design). Using Claude to develop competing models already violates our Terms of Service, but enforcing this restriction through our safeguards avoids accelerating the actors most willing to violate these terms. Unlike our interventions for cybersecurity, biology and chemistry, and distillation attempts, these safeguards will not be visible to the user. Fable 5 will not fall back to a different model. Instead, the safeguards will limit effectiveness through methods such as prompt modification, steering vectors, or parameter-efficient fine-tuning (PEFT). These interventions will not affect the vast majority of coding work. We estimate they will impact ~0.03% of traffic, concentrated in fewer than 0.1% of organizations. I believe this is the first time Anthropic have announced these kinds of silent interventions. The justification still feels pretty science-fiction to me - the linked article talks about "recursive self-improvement". I'm not at all keen on a model that silently corrupts its replies to questions about "ML accelerator design" purely to slow down research that might conflict with Anthropic's own goals! Update: Anthropic walked back this policy in the face of widespread outrage from the research community. Recent articles - Initial impressions of Claude Fable 5 - 9th June 2026 - Running Python code in a sandbox with MicroPython and WASM - 6th June 2026 - Claude Opus 4.8: "a modest but tangible improvement" - 28th May 2026

Items scoring 7/10 or above from 11 sources, scored by claude-haiku-4-5-20251001 on relevance to my interests. At most 3 per source.

Scoring categories & sources
  1. Local Wexford or South East Ireland news
  2. Irish or EU-wide affairs affecting citizens broadly: elections, new laws or policy being debated, cost of living, education — especially impacts on mid-life adults or teenagers. Never courts/crime stories.
  3. Irish news on a topic relevant to my interests
  4. Work and tech topics: networking, AI, Kubernetes, platforms, SaaS
  5. AI news including critical or anti-AI perspectives
  6. Gaming: PC gaming, indie gaming, retro gaming
  7. General interests: gardening, woodwork, cycling, fitness, travel
  8. Comics

Sources: Breaking News Ireland, Wexford Local, Hacker News, r/gaming, r/pcgaming, r/antiAI, r/indiegaming, Lenny's Newsletter, One Useful Thing, Newcomer, Simon Willison

Comics

Beam Pipe

XKCD · view →