daily

2026-07-05
1

80 Co. Wexford projects get funding

Wexford Local · original → · 8/10 · Local Wexford: €250k community funding for 80 projects
By Dan Walsh Minister James Browne TD and his Cabinet Colleague, Minister for Rural and Community Development and the Gaeltacht, Dara Calleary TD, this week announced successful projects under the…

By Dan Walsh

Minister James Browne TD and his Cabinet Colleague, Minister for Rural and Community Development and the Gaeltacht, Dara Calleary TD, this week announced successful projects under the Local Enhancement Programme (LEP) 2026. In Co. Wexford 80 projects will benefit from over €250,000 under the Local Enhancement Programme 2026.

Regarding the Co. Wexford funding, Minister Browne said; “These groups are genuinely so essential to community life as we know it, from groups in Templeshannon and Bunclody to Taghmon, Bridgetown, and New Ross.

“Every single week I meet the people and groups that benefit from grants under the LEP across the county. There are some incredible people giving their time to lead, contribute to and keep these groups going for everyone’s benefit and we are lucky to have them.

“These pots of funding are about improving local facilities through small-scale funding for a wide range of projects including minor renovations to buildings and the purchase of essential equipment for the maintenance of public areas.

“It also allows groups to purchase items such as IT equipment to enable groups to conduct their business.

The 80 groups in Co. Wexford who will get an injection of funding for their activities are in the accompanying list.

[image →]
JAMES BROWNE TD Minister for Housing, Local Government and Heritage. (File Pic; WexfordLocal.com)

12th Wexford Ballykelly Scouts €4,700.

13th Wexford Scout Group (Wexford Town) €4,500.

Access 2000 €5,000.

Adamstown Community Centre €1,589.

Adamstown Community Development Association €5,000.

Ballaghkeen Community Project Ltd €5,000.

Ballycanew Christmas Lights Committee €5,000.

Barnardos (Wexford) €4,505.27

Bellefield Youth Club €2,000.

Bridgetown Men’s Shed €2,882.13

Bunclody Area Tourism €1,000.

Bunclody Bridge Club €870.

Bunclody Tidy Towns €2,000.

Campile Area Development Group €5,000.

Castlebridge Horticultural and Agricultural Show Committee €1,000.

Celtic Dragon KBX Kickboxing €4,925.52

Clonard Girl Guides €1,000.

Community Catholics (Brazilian) in Enniscorthy €1,646.46

Coolgreaney Ballyfad Woods Nature Trails €4,300.

Coolgreany Handball Club €4,800.

Coolgreany Tidy Towns €1,800.

Courtown Communtiy Council €3,030.83

Courtown Heritage Group €2,241.

Courtown Hibs AFC €4,890.

Courtown Tidy Towns €4,604.

Enniscorthy Delightful Dollies Women’s Group €800.

Enniscorthy Basketball Club €491.

Enniscorthy Men’s Shed €744.35

Enniscorthy United FC €5,000.

FAB Coolcotts Development Project CLG €5,000.

FDYS Scoil Spraoi Na Leanai Preschool €2,314.11

FDYS Traveller Inclusion Programme €1,000.

Ferns Community Development Association (FCDA) – Ferns Old Graveyard Steering Group €349.85

Friends Together New Ross- FDYS €1,000.00

Glenbrien Community Hub €5,000.

Gorey Garden City Residents Association €4,600.

Gorey Athletics Club €4,403.

Gorey Boxing Club €2,998.79

Gorey Tidy Towns €1,000.

Kilmore Athletics Club €902.88

Kilmore Quay Community Development Association CLG €5,000.

Kilrane/Rosslare Harbour Men’s Shed €2,595.

Kilrush Brownies €1,000.

New Ross Community Hospital CLG, t/a New Ross Community Care Home €5,000.

Newbawn Development Group €1,000.

Old Kilnenor Historical Society (Graveyard) €4,595.

Pairc Na nGabhar Committee (Bunclody) €5,000.

Poulpeasty Women’s Group €798.

Revive New Ross €5,000.

Riverchapel & Courtown Ladies Club €853.89

Saint Brendan’s Estate Residents €2,020.

Scoil Mhuire Ballyhogue €4,443.

Slaney Search & Rescue €5,000.

Southend Family Resource Centre €4,122.

St Abbans Adamstown GAA Club €4,934.95

St Abbans Kilmyshall Boxing Club €5,000.

St Patrick’s Fife And Drum Band €5,000.

St. Aiden’s Communty Hall (Bunclody) €510.38 St. Columba’s Old Folk’s Club €5,000.

St. John’s Volunteers GAA Club €970.

St. Mary’s Day Care Centre CLG (Tagoat) €1,376.58

St. Mary’s Maudlintown GAA Club €5,000.

Summerhill Girl Guides €1,000.

Sunflower Ukrainian Hub €4,867.94

Taghmon Badminton Club €3,000.

Taghmon Camross GAA Club €3,250.

Taghmon Camross LGFA €837.15

Taghmon Family Resource Centre €2,780.76

Templeshannon Indoor Bowls €2,500.

The May Byrne Community House €2,463.

The Mna le Cheile (Bunclody) €5,000.

The Phoenix Rising Support Network-CLG €1,000.

Vision for Courtown and Riverchapel €4,950.

Wexford Albion FC €4,799.96

Wexford Astro Football €1,898.90

Wexford CBS Boxing Club €5,000.

Wexford Eagles American Football Club €5,000.

Wexford Sub Aqua Club €2,720.

Write By The Sea Literary Festival (WBTS) at Kilmore Quay €4,051.86

Yola Hedge School (Tagoat) €5,000.

2

'We want a candidate in each constituency': TD says 'positive politics' behind Social Democrats' rise

Breaking News Ireland · original → · 7/10 · Irish politics: Social Democrats election strategy and polling
The Social Democrats' popularity is on the rise after a recent Dublin Central byelection victory and opinion poll success, and a senior TD in the party has said they plan to run a candidate in each…

The Social Democrats' popularity is on the rise after a recent Dublin Central byelection victory and opinion poll success, and a senior TD in the party has said they plan to run a candidate in each constituency at the next general election. Daniel Ennis secured a seat in a packed field in Dublin Central, bringing the party's number of sitting TDs to 12. Meanwhile a recent Red C/Business Post poll put the party on 12 per cent, just two points behind Fianna Fáil. Holly Cairns also regularly comes out on top as the most popular party leader. In an interview with BreakingNews.ie, Jennifer Whitmore said: "What we hear from lots of people is they would love to have the chance to vote Social Democrats, but we may not have had a candidate in their area. "We're going to run a candidate in each of the constituencies because we do want to give people that choice. "Yes. Holly has been very clear on this: we want a candidate in each constituency. "We're hoping the next general election will be a good one for us. At each milestone we're hitting we are doing very well, and we would hope that will continue into the next local and general elections." Whitmore, who was first elected in 2020 before retaining her seat in the 2024 general election, said she is confident the party can take votes from people who have become dissatisfied with Fianna Fáil and Fine Gael, as well as opposition rivals like Sinn Féin. 'Positive politics' "I think we will get votes from across the board. What we're offering is something different. I think we're offering positive politics. We really saw that play out in the Dublin Central byelection. Our profile, and how we attract people, might be different from what they're used to; we're looking for positive solutions and to be a positive force in politics. "We have our principles and we stick with them, and I think that's really important because people know what they get with us. That is a vital part of our identity." Whitmore pointed to the shift to the right from the Government and Sinn Féin, and said the fact the Social Democrats have "stuck to our principles" appeals to people. "We have seen a shift to the right in politics, but we have stuck with our principles. I think people value that aspect of who we are." Whitmore said she does not feel rising anti-migrant sentiment represents the average Irish voter, and that people's real frustration is in a shortage of services. "We have our own migration history, so I think as a nation we will be coming at it quite differently. I think the frustration we've seen over the last few years is more about the lack of services. There are some bad actors using that for anti-migration purposes, a bouoying up of frustration as this government and the previous government did not deliver the services people need. "The thing is we are a really wealthy country, we have so much money in the coffers, yet people are not experiencing that on the ground so it does not feel like a wealthy country to people. "People still can't get basic things like GP services, assessments for their children, disability supports, public transport that actually runs on time. "I think a lot of the migration debate is being fuelled by that basic lack of delivery by the Government. For us as a party we will continue to focus on service delivery and the Government's inability to deliver, rather than blaming a cohort of people who may be coming in to work in areas we need to shore up like the health service or in childcare. We're not going to point the finger at those people, we will point it at Government because people do not have the services they need." She added: "We're in the business of politics to make people's lives better and I think you can feel with other parties sometimes their focus is more on who will be the next leader, who's saying what, all these internal political bubble things that take over their discussions. "Most people don't want to see that or hear about it... they just want to see a government and a Dáil that is working for them and delivering what people need. That is the Social Democrats' focus. "There is a cost of living crisis... what can we do to push the Government to support people to deal with that crisis? "Our focus is on these issues, what is needed for communities, families and small businesses." When Catherine Murphy and Róisín Shortall stepped down as co-leaders of the party in 2023, there was no leadership contest. Whitmore said this points to how the Social Democrats are different, with the party's TDs working as a unit. She said this was behind the "seamless" transition when deputy leader Cian O'Callaghan stepped in during Cairns' maternity leave. Whitmore is the Social Democrats spokesperson on climate, energy and biodiversity. She described the Government's current position on climate action as "worrying". "They have gone back on a lot of climate action. I think one of the failures of the last government is climate action was seen as something that was punitive, whereas actually if climate action is done properly it can be a positive thing for people. "If people are given proper support on climate action you're talking warmer homes, cheaper electricity bills, cleaner air to breathe. These are positive things that result from a proper approach to climate action. "I think it is worrying the Government are rowing back on this, probably as a result of the previous government not getting that message out. You need to meet people where they're at, and at the moment people are really struggling with energy and food costs. "You can provide affordable energy in a way to help with climate targets, the focus we need is on how we get cheap, renewable energy to people." Whitmore feels a left wing government is possible, despite the obvious challenges of such a coalition. She again reiterated the positivity around her party, adding that she feels this marks them out from their rivals. "Our preference is for a left wing government and we think that will deliver the best for people... but we have to be open to talking to everybody. "We want to be as big as possible as a party entering those negotiations. "There's a real possibility for us. People are open to our ideals and what we're presenting to them and I would hope we can grow as much as possible. "Politics can feel very draining, and when people are bombarded with negativity they will disengage. People want to hear things can be fixed, can get better, and we are a party willing to fight for those solutions for people."

3

Command and Conquer Generals natively ported to macOS, iPhone, iPad using Fable

Hacker News · original → · 7/10 · Gaming: retro game ported to modern platforms
Zero Hour running natively on Apple Silicon Macs, iPhone, and iPad — campaign, skirmish, and Generals Challenge, with touch controls built for RTS (tap-select, drag-box, long-press deselect,…

Zero Hour running natively on Apple Silicon Macs, iPhone, and iPad — campaign, skirmish, and Generals Challenge, with touch controls built for RTS (tap-select, drag-box, long-press deselect, two-finger scroll, pinch zoom). No emulation: this is the real 2003 engine compiled for ARM64, rendering DirectX 8 → DXVK → Vulkan → MoltenVK → Metal. Built on EA's GPL v3 source release via fbraz3/GeneralsX (which did the heavy lifting of the macOS/Linux port — this fork adds the iOS/iPadOS port and a set of engine fixes). The original GeneralsX README lives on the upstream-main branch. No game assets are included or distributed. You need your own copy (Steam, ~$5 on sale). Prerequisites (one time): # Toolchain xcode-select --install brew install cmake ninja meson pkgconf brew install --cask steamcmd # vcpkg (full clone — a shallow clone breaks manifest baselines) git clone https://github.com/microsoft/vcpkg ~/vcpkg && ~/vcpkg/bootstrap-vcpkg.sh export VCPKG_ROOT=~/vcpkg # add to your shell profile # LunarG Vulkan SDK (NOT the Homebrew cask) — https://vulkan.lunarg.com/sdk/home export VULKAN_SDK=$HOME/VulkanSDK/<version>/macOS # add to your shell profile Clone, build, get assets, play: git clone https://github.com/ammaarreshi/Generals-Mac-iOS-iPad.git GeneralsX cd GeneralsX ./scripts/build/macos/build-macos-zh.sh # checks deps, configures, builds ./scripts/build/macos/deploy-macos-zh.sh # creates ~/GeneralsX/GeneralsZH + run.sh ./scripts/get-assets.sh <your_steam_username> # fetches game data you own cd ~/GeneralsX/GeneralsZH && ./run.sh -win On top of the macOS prerequisites: full Xcode (signed into your Apple ID), brew install xcodegen , and a (free or paid) Apple Developer team. cd GeneralsX git submodule update --init references/fbraz3-dxvk # iOS DXVK is built from this + Patches/dxvk-ios.patch ./scripts/build/ios/fetch-moltenvk.sh # pinned MoltenVK.framework (checksummed) ./scripts/build/ios/stage-fonts.sh # Liberation fonts, renamed as the game expects cmake --preset ios-vulkan cmake --build build/ios-vulkan --target z_generals GX_TEAM_ID=<your-team-id> GX_BUNDLE_ID=com.you.generalszh \ ./scripts/build/ios/package-ios-zh.sh --install # assembles, signs, installs Find your team id in Xcode → Settings → Accounts. Assets ship inside the app bundle (self-contained install); --dev skips the ~2.7 GB copy for fast code iteration. | Path | What it is | |---|---| docs/port/PORTING_PLAYBOOK.md | The complete engineering log of this port: every failure mode, root cause, fix — start with §8, the bug archaeology: the black minimap, the silent EVA lines, and the chirp | docs/port/PORTING_PATTERNS.md | Generalized methodology for porting classic Windows games to Apple platforms | docs/port/RELEASE_CHECKLIST.md | Gate for public release | scripts/get-assets.sh | Steam asset fetcher (your own copy; app 2732960) | scripts/build/macos/ , scripts/build/ios/ | Build, deploy, packaging pipelines | ios/ | XcodeGen signing-stub project + ios/config/ (staged Options.ini, dxvk.conf) | Patches/dxvk-ios.patch | DXVK changes the iOS d3d8/d3d9 dylibs are built from (applied via the local-fork build) | - Long sessions on iPad can be killed by iOS for memory (~3 GB+ resident); the app exits to the home screen with no dialog. Session logs (current + previous) are in the Files app under the game's folder. Under investigation. - Backgrounding mid-game can occasionally crash on iOS — the lifecycle pause covers the common paths; a rare race remains. Save often. Engine code GPL v3 (EA's source release → GeneralsX → this fork). Game assets: not included, not licensed here. Credits: Westwood/EA Pacific (the game), EA (the source release), fbraz3/GeneralsX (the base port), TheSuperHackers/GeneralsGameCode (community mainline), DXVK, MoltenVK, SDL, OpenAL Soft, FFmpeg, Liberation Fonts. This port was built as a human+AI collaboration: engineering by Claude Code (Anthropic's Claude, Fable model), directed and playtested on real devices by Ammaar Reshi. The engineering log in docs/port/ is the unedited record of how that worked.

4

The Log Is the Agent

Hacker News · original → · 7/10 · AI: event-sourced AI agent architecture paper
Computer Science > Artificial Intelligence [Submitted on 21 May 2026] Title:The Log is the Agent: Event-Sourced Reactive Graphs for Auditable, Forkable Agentic Systems View PDF HTML…

Computer Science > Artificial Intelligence [Submitted on 21 May 2026] Title:The Log is the Agent: Event-Sourced Reactive Graphs for Auditable, Forkable Agentic Systems View PDF HTML (experimental)Abstract:Most agent frameworks are built around the language model: a conversation loop comes first, then tools, then rules, and finally a logging layer bolted on for observability, with state persisted as retrievable "memory." We describe ActiveGraph, a runtime that inverts this arrangement. The append-only event log is the source of truth; the working graph is a deterministic projection of that log; and behaviors--ordinary functions, classes, LLM-backed routines, or logic attached to typed edges--react to changes in the graph and emit new events. No component instructs another; coordination happens entirely through the shared graph. This single design decision yields three properties that retrieval-and-summarization memory systems do not provide: deterministic replay of any run from its log, cheap forking that branches a run at any event without re-executing the shared prefix, and end-to-end lineage from a high-level goal down to the individual model call that produced each artifact. We present the architecture, a determinism contract that makes replay sound, and a worked diligence example whose full causal structure is reconstructable from the log alone. We discuss--without claiming to demonstrate--why this substrate is unusually well suited to self-improving agents, and how it extends the BabyAGI lineage and prior graph-memory research. References & Citations Loading... Bibliographic and Citation Tools Bibliographic Explorer (What is the Explorer?) Connected Papers (What is Connected Papers?) Litmaps (What is Litmaps?) scite Smart Citations (What are Smart Citations?) Code, Data and Media Associated with this Article alphaXiv (What is alphaXiv?) CatalyzeX Code Finder for Papers (What is CatalyzeX?) DagsHub (What is DagsHub?) Gotit.pub (What is GotitPub?) Hugging Face (What is Huggingface?) ScienceCast (What is ScienceCast?) Demos Recommenders and Search Tools Influence Flower (What are Influence Flowers?) CORE Recommender (What is CORE?) arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.

5

Better Models: Worse Tools

Hacker News · original → · 7/10 · AI: Claude model tool-calling degradation analysis
written on July 04, 2026 A very strange Pi issue sent me down a rabbit hole over the last two days. The short version is that newer Claude models sometimes call Pi’s edit tool with extra, invented…

written on July 04, 2026 A very strange Pi issue sent me down a rabbit hole over the last two days. The short version is that newer Claude models sometimes call Pi’s edit tool with extra, invented fields in the nested edits[] array. And not Haiku or some small model: Opus 4.8. The edit itself is usually correct but the arguments do not match the schema as the model invents made-up keys and Pi thus rejects the tool call and asks to try again. That alone is not too surprising as models emit malformed tool calls sometimes. Particularly small ones. What surprised me is that this is getting worse with newer Anthropic models as both Opus 4.8 and Sonnet 5 show it but none of the older models. In other words, the SOTA models of the family are worse at this specific tool schema than their older siblings. In case you are curious about Fable: I intentionally did not test it because I was not sure if the classifiers they are running might downgrade me to Opus silently. If you have not spent too much time looking at LLM tool calling internals, the important thing to understand is that tool calls are not magic and use some rather crude in-band signalling. The model receives a transcript, a system prompt and a list of available tools. The server munches that into a large prompt with special marker tokens. Because the model was trained and reinforced on examples of that format, at some point during generation it emits something that the API or client interprets as “call this tool with these arguments”. For a file edit tool, the intended invocation payload might say something like this: { "path": "some/file.py", "edits": [ { "oldText": "text to replace", "newText": "replacement text" } ] } A harness then validates the arguments, performs the edit, and feeds the result back into the model. If validation fails, the model sees an error and usually tries again. How exactly that formatting happens is not known for the Anthropic models, but some people have gotten out “ANTML” markers and they at times do leak also into public communications. To the best of my knowledge, the call above would come out serialized like this from the model: <antml:function_calls> <antml:invoke name="edit"> <antml:parameter name="path">some/file.py</antml:parameter> <antml:parameter name="edits"> [ { "oldText": "text to replace", "newText": "replacement text" } ] </antml:parameter> </antml:invoke> </antml:function_calls> An important thing to note here is that this thing, while looking like XML, is not really XML. It’s just a thing they found convenient to tokenize and train on. The other thing to note is that a basic top-level string parameter appears in-line whereas an array of objects is implemented via JSON serialization. While I’m not entirely sure that this is how it works, there are some indications that this is not too far off. This will become relevant later. There are two very different ways to make the model produce a structure like this: The second approach is what people usually refer to as grammar-aware or constrained decoding. The sampler masks out tokens that would violate the grammar. If the model is currently inside a JSON object and the schema says only oldText and newText are allowed, the sampler can prevent it from emitting "in_file" or "type" . Grammar-aware decoding can be used both to constrain something to be syntactically valid JSON and also to enforce specific enum values or keys. Without any form of constraints the model is merely following a learned convention. Pi’s edit tool supports multiple exact string replacements in one call. That is why the arguments contain an edits array. In the failing cases the model produces entries like this: { "oldText": "...", "newText": "...", "requireUnique": true } or this: { "oldText": "...", "newText": "...", "oldText2": "", "newText2": "" } Across repeated trials I saw a whole zoo of invented trailing keys: type , id , kind , unique , requireUnique , matchCase , in_file , forceMatchCount , children , notes , cost , oldText2 , newText2 , oldText_2 , newText_2 , and even an event.0.additionalProperties key inside the edit object itself. The most annoying part is that the actual oldText and newText payloads were byte-correct in the invalid calls I inspected. The model had in fact produced the right invocation but then added nonsense at the end of the object. The failure is also heavily context-dependent. A fresh single-turn prompt like “edit this file” did not reproduce it at all for me. An agentic history where the model had read files, diagnosed a problem and then composed a multi-line edit could reproduce it. And more annoyingly, not all transcripts will show that behavior. In fact, I needed Petr Baudis‘s transcripts to reproduce this for me at all! In that user’s session continuing the session caused Opus 4.8 to fail around 20% of the time. Stripping thinking blocks from history reduced the failure rate by half. Turning on strict tool invocation eliminated it in my runs. My strongest hypothesis is that this is not random deterioration but a training artifact. When older Anthropic models were trained, they were trained on some tools (some of which were documented). But that training did not yet have a user-shipped harness like Claude Code as the obvious target. Modern Anthropic models are most likely different because their post-training includes Claude Code or a harness that looks very similar. The model learns what a successful tool call looks like in that environment. It also learns what mistakes are tolerated by that environment. Claude Code’s own tools are comparatively flat. The ordinary edit tool is not Pi’s nested edits[] shape; it is closer to file_path , old_string , new_string , and an optional flag (replace_all ). Looking at Claude Code’s client is very instructive: it contains retry paths for malformed tool use, parameter aliases, type coercions, Unicode repairs and filtering of unknown keys. In other words, Anthropic’s own client appears to expect and accept a fair amount of slop and repairs it, mostly silently. If reinforcement learning happens in a harness like that, or a simulation of one, then slightly malformed tool calls can still complete the task and receive reward. The harness fully absorbs the error and there is little gradient against inventing an alias, adding a stray field or using a nearby parameter name. Worse, the model may become very strongly adapted to the canonical Claude Code edit tool shape. A different harness can present a tool with the same semantic intent but a different schema. Such a tool can increasingly be off-distribution. The better-trained model might actually fight you harder because its prior is stronger. This is not too surprising, but it is a change from how this was a few months ago. When Opus 4.5 launched, it adapted to other edit tools exceptionally well. In fact, I was pretty convinced that we’re on a good path where the models are more likely to adapt to any sort of tool shape that comes around for as long as the instructions are good. Now I’m somewhat worried about the track we’re on here. Alternative tool schemas might not just be unfamiliar. They might be implicitly punished by post-training that optimizes for one particular, forgiving tool ecology. And that ecology is not documented. While there is a text editor tool that is documented, you will see that this format is in fact not followed by Claude Code. What Claude Code does internally (which is a closed-source harness) is hidden from you. Claude Code is obviously closed-source but we can look at the minified code and get some idea of what it does. And honestly, it’s very forgiving of incoming data. For a start, Claude Code checks the model’s visible text for leaked <invoke markup. It also emits some telemetry when that happens and then it has its own state machine to retry such bad calls by pushing back to the model. It has explicit Unicode escape repair which fixes broken \uXXXX sequences and lone surrogates in string values. It also has per-tool aliases for parameters. For instance, Edit accepts old_str (presumably from the times when the models were trained on the officially documented text editor tool), the newer old_string from the schema, new_str /new_string , path as an alias for file_path , and some more. It also silently filters out unexpected keys and it does not use strict mode either. The issue with strict mode is that Anthropic applies complexity limits to the tool definitions that cause API requests to fail, so presumably that’s why Claude Code does not attempt to use it. Will this problem be with us in other harnesses too? One huge issue with Anthropic is that the models are completely closed, and so is the harness. Codex models are also closed, but at least the harness is not. We also have gpt-oss which is at least a bit interesting. The models are explicitly trained to use OpenAI’s harmony response format and there is a lot of documentation that at least tells us how OpenAI people think about this. Harmony makes channels and tool-call content types part of the prompt format. A function call can look like this: <|start|>assistant<|channel|>commentary to=functions.get_weather <|constrain|>json<|message|>{"location":"San Francisco"}<|call|> The important bit is <|constrain|>json . The model can express in-band that this message body is JSON, and an inference stack can use that boundary to switch into JSON-constrained sampling for the body of the tool call. Presumably a bit of this also happens in Anthropic’s models, at least in strict mode I would imagine. The marker in harmony helps the sampler to detect when it needs to sample with a specific grammar, and because it is part of the transcript, it makes that rather easy to do. For hosted GPT models, there is also an option to provide a LARK grammar for custom tools that need to adhere to something like this. Anthropic appears different from that, though maybe not entirely. If an array of objects is represented as JSON, as it appears to be, then the model has to write JSON inside the tool parameter. There is probably basic grammar-constrained sampling going on, and that may partly explain the extra keys. For a nested array parameter, that JSON includes escaped multi-line file content inside string literals, inside one tag. The unexpected, made-up keys appear exactly at the highest-entropy point of that task: after closing a several-hundred-token escaped newText string, where the model must decide } vs , "..." . Opus 4.8 and Sonnet 5 seem to have much stronger priors about what an edit tool call should look like and that prior appears to be Claude Code’s edit schema: a flat old/new string pair, plus the optional replace_all flag. My guess is that Opus has learned that an edit operation may have one extra optional field, but under Pi’s nested oldText /newText shape it has no trained name for that field. So it samples a plausible name fresh each time, which is why the failures produce dozens of random keys rather than one stable alias. As strict mode in Anthropic appears to fix this, I presume that on the server side they are refusing to sample a key that is not permitted by the JSON schema structure. That would also explain why they have limits to the complexity of the tool definitions when strict mode is enabled. So far, the Codex models I tested did not show this type of regression. I tested all available ones except 5.6, which I do not have access to yet. The uncomfortable lesson is that tool schemas are not neutral, at least not on Anthropic models. We like to pretend that a schema is an abstract contract and the model is a general reasoner that will follow it, but that might no longer be the case for some of the tools. Tool schemas are somewhere in the distribution and some shapes are close to what the model saw during post-training and some are far away. Some are easy for the provider’s hidden encoding (e.g. top-level attributes in ANTML), whereas some require the model to write large escaped JSON objects inside nested arrays after long multiline strings. The model may be smart enough to understand the schema and still be bad at sampling the exact shape under pressure. If this type of model behavior continues, I wonder what the implications for harnesses are. Obviously one could turn on strict sampling in Anthropic and the problem should go away. On the other hand, that the model has this behavior shows the impact that reinforcement learning has on them. Fighting that prior is probably futile if you want to get the best model performance. Right now the reality is that Claude Code is not open source and we cannot really know what they are doing in their RL environments either. We cannot assume Claude-Code-trained behavior will transfer cleanly to your tools unless they are a close match. The more post-training happens inside one dominant harness, the more every other harness will have to inherit its quirks. I used to be more skeptical of strict grammar-constrained tool invocation because constrained decoding can have quality tradeoffs. I still think that can be true in general, but this bug moved my priors significantly. If the newest models get better at solving the task while getting worse at faithfully emitting an alternative tool schema, then the harness needs stronger guarantees somewhere. If you want to find out more, or you want to discuss this, consider reading the issue on the Pi tracker.

6

sqlite-utils 4.0rc2, mostly written by Claude Fable (for about $149.25)

Simon Willison · original → · 7/10 · AI: Claude Fable writing software release notes
sqlite-utils 4.0rc2, mostly written by Claude Fable (for about $149.25) 5th July 2026 I wrote about the sqlite-utils 4.0rc1 release a couple of weeks ago. Since we only have Claude Fable on our Max…

sqlite-utils 4.0rc2, mostly written by Claude Fable (for about $149.25) 5th July 2026 I wrote about the sqlite-utils 4.0rc1 release a couple of weeks ago. Since we only have Claude Fable on our Max subscriptions for a few more days, I decided to see if it could help me get to a 4.0 stable release that I felt truly comfortable about, since I try to keep to SemVer and like my incompatible major versions to be as rare as possible. I started with this prompt, in Claude Code for web on my iPhone: Final review before shipping a stable 4.0 release - very important to spot any last minute things that would be a breaking change if we fix them later Here’s that initial report it created for me. There were some significant problems that I hadn’t myself encountered yet—5 that Fable categorized as “release blockers”. Here’s the worst of the bunch: 1. delete_where() never commits and poisons the connection (data loss) Table.delete_where() (sqlite_utils/db.py:2948 ) runs its DELETE via a bareself.db.execute() with noatomic() wrapper — compareTable.delete() atdb.py:2944 , which wraps correctly. The connection is leftin_transaction=True , so every subsequentatomic() call takes the savepoint branch (db.py:430-440 ) and never commits either.Reproduced end-to-end: db = sqlite_utils.Database("dw.db") db["t"].insert_all([{"id": i} for i in range(3)], pk="id") db["t"].delete_where("id = ?", [0]) # conn.in_transaction is now True db["t"].insert({"id": 50}) db["u"].insert({"a": 1}) db.close() # Reopen: rows are [0, 1, 2] — the delete, row 50, AND table u are all gone. That’s a really bad bug! Very glad I didn’t ship that, although at least it would have been a bug I could fix in a 4.0.1 point release, not a design flaw that would force a 5.0. Over the course of 37 prompts, 34 commits and +1,321 -190 code changes over 30 separate files, we worked through the entire set of feedback in turn, making several other design improvements along the way. A weird thing about coding agents is that harder tasks like this one actually provide more opportunity to do other things at the same time, since the agent sometimes needs 10-15 minutes to churn away on a new task. I went out to enjoy the Half Moon Bay 4th of July parade, occasionally checking in and prompting the next step for Fable from my phone. Full details in the PR and this shared transcript. I switched to my laptop for the final review, which I conducted through GitHub’s PR interface. The most significant changes relate to transaction handling, which was the signature new feature in the earlier RC. The new RC now includes comprehensive documentation on the new transaction model, the intro to which I’ll quote here in full: Every method in this library that writes to the database— insert() ,upsert() ,update() ,delete() ,delete_where() ,transform() ,create_table() ,create_index() ,enable_fts() and the rest—runs inside its own transaction and commits it before returning. Your changes are saved to disk as soon as the method call finishes:db = Database("data.db") db.table("news").insert({"headline": "Dog wins award"}) # The new row is already saved - no commit() requiredThe same applies to raw SQL executed with db.execute()—a write statement is committed as soon as it has run. You never need to call commit() , and you do not need to close the database to persist your changes. There are exactly two situations where you need to think about transactions: You want to group several write operations together, so they either all succeed or all fail—use db.atomic(). You are managing a transaction yourself with db.begin() , in which case nothing is committed until you commit—the library will never commit a transaction you opened. In reviewing Fable’s documentation—I find that reviewing the documentation edits first is an excellent way to build an initial understanding of what has changed—I spotted this detail: db.atomic() and the automatic per-method transactions are designed for connections in Python’s default transaction handling mode. Connections created with the Python 3.12+sqlite3.connect(..., autocommit=True) orautocommit=False options are not supported, becausecommit() androllback() behave differently on those connections. I admit I hadn’t thought about how sqlite-utils would react to the more recent autocommit setting, added in Python 3.12. It turns out “behave differently on those connections” equated to almost the entire test suite failing, so I worked with the model to ensure that this difference would not break how the library works. And a final review by GPT-5.5 I used to think that the idea of having one model review the work of another was somewhat absurd—it felt weirdly superstitious. The problem is it really does work—I’ve started habitually having Anthropic’s best model review OpenAI’s work and vice versa, because I’ve had that turn up interesting results often enough to be valuable. I prompted Codex Desktop and GPT-5.5 xhigh with the following: Review changes since the last RC. Also confirm that the changelog is up-to-date. Which was enough to turn up two issues worth investigating: Findings - [P1] sqlite_utils/db.py:663 db.query() now rejects non-row statements only after callingdb.execute() , and sqlite_utils/db.py:705 auto-commits those writes first. Sodb.query("update ...") raisesValueError but the update is already committed. That is a surprising side effect for a method documented as “can only be used with SQL that returns rows.”- [P1] sqlite_utils/db.py:672 INSERT ... RETURNING throughdb.query() only commits after the returned generator is fully exhausted.db.query("insert ... returning ...") without iteration, or commonnext(db.query(...)) usage, leaves the transaction open and the write can be rolled back on close. This contradicts docs/changelog.rst:15 and docs/python-api.rst:232, which say it takes effect without iteration. I pasted that into a fresh Fable session, which ran some experiments to confirm the problem: Both findings were confirmed. db.query() calledself.execute() first, which auto-commits writes, and only then checkedcursor.description — sodb.query("update ...") committed the update before raisingValueError . And theINSERT ... RETURNING commit lived at the end of the returned generator, so it never fired unless you exhausted the iterator —next(db.query(...)) or an un-iterated call left the transaction open, contradicting what the changelog and docs promise. Here’s the PR with the fix, and the full Claude Code transcript. Reviewing this code helped me build a better mental model of the edge cases of SQLite transaction semantics! For an estimated (unsubsidized) cost of $149.25 I upgraded to the Claude Max $200/month plan (I was previously on $100/month) to increase my Fable allowance for the remaining time until the July 7th Fablepocalypse, when even Claude Max subscribers will have to pay full API cost for the model. I was curious as to how much this would have cost me if I had been paying those costs directly. At first I thought those numbers weren’t available to me since I had run the work remotely using Claude Code for web, and then I realized I could run AgentsView inside that existing session to get that cost estimate! Run "uvx agentsview --help" and then use that tool to calculate the cost of this session Claude figured out how to use the session list --include-children command and came out with the following: | Transcript | Model | Cost | |---|---|---| | Main session | claude-fable-5 | $141.02 | | API-surface sweep agent | claude-fable-5 | $2.40 | | Transactions/atomic review agent | claude-fable-5 | $2.39 | | Post-rc1 commits review agent | claude-fable-5 | $1.72 | | Migrations review agent | claude-fable-5 | $1.40 | | Prompt-counting agent | claude-opus-4-8 | $0.32 | | Total | $149.25 | I’m very glad I’m on that subscription! I really should have followed my own advice and leaned more heavily into subagents with cheaper models. Here’s what claude.ai/settings/usage is showing me right now: I have several other major Fable-driven projects on the go right now as well, with the goal of hitting 100% on that Fable bar just in time for the price increase. The full release notes for sqlite-utils 4.0rc2 Here are the full release notes for the RC. I had Fable add these to an “Unreleased” section of the changelog as each change landed, reviewing them as it went. This has the neat side effect that the commit history of the changelog acts as a concise summary of each of the changes that went into the release. In the past I’ve had a policy of writing release notes by hand, but honestly these are better than I would have created myself. Release notes are a great example of writing that I’m OK to outsource to agents because they need to be boring, predictable and accurate. Breaking changes: - Write statements executed with db.execute() are now committed automatically, unless a transaction is already open in which case they join it. Previously they opened an implicit transaction that stayed open until something committed it—writes appeared to work when read on the same connection but were silently rolled back when the connection closed. Code that relied on rolling back uncommitteddb.execute() writes should use the newdb.begin() method to open an explicit transaction first. The transaction model is documented in full at Transactions and saving your changes.db.query() now executes its SQL as soon as it is called, rather than waiting until the returned generator is first iterated. Rows are still fetched lazily during iteration. SQL errors are now raised at the call site, statements such asINSERT ... RETURNING are executed and committed immediately without needing to iterate over their results, and passing a statement that returns no rows—previously a silent no-op—now raises aValueError recommendingdb.execute() instead. A statement rejected this way is rolled back before the error is raised, so it has no effect on the database.- Python API validation errors now raise ValueError instead ofAssertionError . Previously invalid arguments—such ascreate_table() with no columns,transform() on a table that does not exist, or passing bothignore=True andreplace=True —were rejected using bareassert statements, which are silently skipped when Python runs with the-O flag. Code that caughtAssertionError for these cases should catchValueError instead.table.upsert() andtable.upsert_all() now raisePrimaryKeyRequired if a record is missing a value for any primary key column, or has a value ofNone for one. Previously such records—which can never match an existing row—were quietly inserted as brand new rows, or triggered a confusingKeyError after the insert had already taken place.db.enable_wal() anddb.disable_wal() now raise asqlite_utils.db.TransactionError if called while a transaction is open. Previously they would silently commit the open transaction as a side effect of changing the journal mode, breaking the rollback guarantee ofdb.atomic() and of user-managed transactions.- The View class no longer has anenable_fts() method. It existed only to raiseNotImplementedError , since full-text search is not supported for views—calling it now raisesAttributeError instead, and the method no longer appears in the API reference. Thesqlite-utils enable-fts command shows a clean error when pointed at a view.- The no-op -d/--detect-types flag has been removed from theinsert andupsert commands. Type detection has been the default for CSV/TSV data since 4.0a1, so the flag did nothing—invocations using it should simply drop it.--no-detect-types remains available to disable detection.Database() now raises asqlite_utils.db.TransactionError if passed a connection created with the Python 3.12+sqlite3.connect(..., autocommit=True) orautocommit=False options.commit() androllback() behave differently on those connections, which previously caused every write made by the library to be silently discarded when the connection closed.Everything else: - Fixed a bug where table.delete_where() ,table.optimize() andtable.rebuild_fts() did not commit their changes, leaving the connection inside an open transaction. Their work—and any subsequent writes—could then be silently rolled back when the connection was closed. All three now usedb.atomic() , consistent with the other write methods.- The sqlite-utils drop-table command now refuses to drop a view, anddrop-view refuses to drop a table. Previously each would silently drop the wrong type of object if the name matched. Both now exit with an error suggesting the correct command to use.- Migrations applied by the new migrations system now run inside a transaction, together with the record of the migration having been applied. If a migration raises an exception its changes are rolled back and it stays pending, so it can be safely re-applied after the error is fixed. Migrations that cannot run inside a transaction, such as those executing VACUUM , can opt out using@migrations(transactional=False) —see Migrations and transactions.table.upsert() andtable.upsert_all() now detect the primary key or compound primary key of an existing table, so thepk= argument is no longer required when upserting into a table that already has a primary key.db.table(table_name).insert({}) can now be used to insert a row consisting entirely of default values into an existing table, usingINSERT INTO ... DEFAULT VALUES . (#759)- Improvements to the sqlite-utils migrate command:--stop-before values that do not match any known migration are now an error instead of being silently ignored,--stop-before now works correctly with migration files that still use the oldersqlite_migrate.Migrations class, and--list is now a read-only operation that no longer creates the database file or the migrations tracking table.migrations.applied() now returns migrations in the order they were applied.- New db.begin() ,db.commit() anddb.rollback() methods for taking manual control of transactions, as an alternative to thedb.atomic() context manager.- New documentation: Transactions and saving your changes describes how transactions work and when changes are committed, and a new Upgrading page details the changes needed to move between major versions.

7

Better Models: Worse Tools

Simon Willison · original → · 7/10 · AI: Claude model tool-calling issues analysis
4th July 2026 - Link Blog Better Models: Worse Tools. Armin reports on a weird problem he ran into while hacking on Pi: The short version is that newer Claude models sometimes call Pi’s edit tool…

4th July 2026 - Link Blog Better Models: Worse Tools. Armin reports on a weird problem he ran into while hacking on Pi: The short version is that newer Claude models sometimes call Pi’s edit tool with extra, invented fields in the nested edits[] array. And not Haiku or some small model: Opus 4.8. The edit itself is usually correct but the arguments do not match the schema as the model invents made-up keys and Pi thus rejects the tool call and asks to try again.That alone is not too surprising as models emit malformed tool calls sometimes. Particularly small ones. What surprised me is that this is getting worse with newer Anthropic models as both Opus 4.8 and Sonnet 5 show it but none of the older models. In other words, the SOTA models of the family are worse at this specific tool schema than their older siblings. Armin theorizes that this is because more recent Anthropic models have been specifically trained (presumably via Reinforcement Learning) to better use the edit tools that are baked into Claude Code. This has the unfortunate effect that other coding harnesses, such as Pi, may find that their own custom edit tools are more likely to be used incorrectly. Claude's edit tool uses search and replace. OpenAI's Codex uses an apply_patch mechanism instead, and OpenAI have talked in the past about how their models are trained to use that tool effectively. Does this mean third-party coding harnesses like Pi should implement multiple edit tools just so they can use the one with the best performance for the underlying model the user has selected?

Items scoring 7/10 or above from 11 sources, scored by claude-haiku-4-5-20251001 on relevance to my interests. At most 3 per source.

Scoring categories & sources
  1. Local Wexford or South East Ireland news
  2. Irish or EU-wide affairs affecting citizens broadly: elections, new laws or policy being debated, cost of living, education — especially impacts on mid-life adults or teenagers. Never courts/crime stories.
  3. Irish news on a topic relevant to my interests
  4. Work and tech topics: networking, AI, Kubernetes, platforms, SaaS
  5. AI news including critical or anti-AI perspectives
  6. Gaming: PC gaming, indie gaming, retro gaming
  7. General interests: gardening, woodwork, cycling, fitness, travel
  8. Comics

Sources: Breaking News Ireland, Wexford Local, Hacker News, r/gaming, r/pcgaming, r/antiAI, r/indiegaming, Lenny's Newsletter, One Useful Thing, Newcomer, Simon Willison