daily

2026-09-06
1

Introducing GPT-6 Astra for developers

Simon Willison · original → · 8/10 · AI: GPT-6 Astra developer capabilities and 3D modeling
5th September 2026 - Link Blog Introducing GPT-6 Astra for developers (via) Blink and you'll miss it, but there's a familiar creature at 1m59s: Across the board, Astra has more attention to detail,…

5th September 2026 - Link Blog Introducing GPT-6 Astra for developers (via) Blink and you'll miss it, but there's a familiar creature at 1m59s: Across the board, Astra has more attention to detail, better understanding of the user's prompt, and can build more sophisticated outputs. In particular, it excels at building 3D models. I've seen it make incredible renderings of gardens, shipyards, animals, cityscapes, even Dyson spheres. Astra really does believe in putting a red neckerchief on a pelican riding a bicycle. Recent articles - The Pelican comparison grid for Astra is pretty interesting - 4th September 2026 - OpenAI's rogue agents were caught communicating via public wikis - 4th September 2026 - Claude's new system prompt really doesn't want to reproduce song lyrics - 2nd September 2026

2

GPT-6 Astra on robot arms

Hacker News · original → · 7/10 · AI: GPT-6 Astra robotic manipulation capabilities, practical AI application
GPT‑6 Astra on robotic manipulation September 4, 2026 A follow‑up to our comparison of Claude Fable 5 and Fable 5.1. We gave OpenAI's GPT‑6 Astra control of the same YAM arms under the same Inspect…

GPT‑6 Astra on robotic manipulation September 4, 2026 A follow‑up to our comparison of Claude Fable 5 and Fable 5.1. We gave OpenAI's GPT‑6 Astra control of the same YAM arms under the same Inspect Robots agent policy, on the same two tasks: “Pick up the red block from the table and place it inside the bowl.” “Pick up the round blue puzzle piece by the knob at its center and place it into the matching circular groove in the board.” On the bowl task Astra placed the block in 19 of 20 trials, against Fable 5.1's 8 of 20 and Fable 5 in 1 of 20, in 2.5 minutes per trial to Fable 5.1's 6.8, at an estimated $0.94 per run to $2.12. The puzzle task is a different story: Astra completed the insertion 2 times in 20 against Fable 5.1's 2 in 20. It reaches the groove and stalls at the same final step Fable does, at $1.36 per run to $2.18. Block into bowl: the best completed run of each model (highest stage, then shortest), each played in its own time at the same speed‑up. Timers show real elapsed time with thinking pauses removed. Astra completes the bowl task far more often, at about half the cost per run Large dots are condition means; faint dots are individual trials (100 if completed, 0 otherwise) at their own cost. Scoring Every trial was scored by a human grader on the highest stage it reached, so a run that fails still records how far it got. The rubric is unchanged from the Fable report. | 0 | No purposeful approach | | 1 | Made contact with the object | | 2 | Lifted the object clear of the table | | 3 | Positioned it above the deposit point | | 4 | Placed it in its final position | Astra places the block almost every time; on the puzzle it stalls where Fable does Share of trials per model reaching each stage; n per row is the number of trials in that cell. Results | Task | Model | Mean stage | Completions | Rate | Output tokens/run | Est. cost/run | Minutes/run | |---|---|---|---|---|---|---|---| | Block into bowl | Fable 5 | 1.30 | 1 / 20 | 5% | 19.2k | $2.69 | 8.2 | | Block into bowl | Fable 5.1 | 2.40 | 8 / 20 | 40% | 12.9k | $2.12 | 6.8 | | Block into bowl | GPT-6 Astra | 3.95 | 19 / 20 | 95% | 2.1k | $0.94 | 2.5 | | Puzzle into groove | Fable 5 | 1.50 | 0 / 20 | 0% | 16.3k | $2.63 | 7.9 | | Puzzle into groove | Fable 5.1 | 2.35 | 2 / 20 | 10% | 10.5k | $2.18 | 5.9 | | Puzzle into groove | GPT-6 Astra | 2.00 | 2 / 20 | 10% | 2.7k | $1.36 | 3.4 | All runs Every counted trial, 120 in total. | Task | Model | Score | Time (min) | Output tokens | Transcript | Video | Rerun | |---|---|---|---|---|---|---|---| | Block into bowl | GPT-6 Astra | 4 | 1.7 | 1,533 | view | watch | download | | Block into bowl | GPT-6 Astra | 4 | 1.7 | 1,734 | view | watch | download | | Block into bowl | GPT-6 Astra | 4 | 1.7 | 1,573 | view | watch | download | | Block into bowl | GPT-6 Astra | 4 | 1.8 | 1,669 | view | watch | download | | Block into bowl | GPT-6 Astra | 4 | 1.9 | 1,570 | view | watch | download | | Block into bowl | GPT-6 Astra | 4 | 2.1 | 1,491 | view | watch | download | | Block into bowl | GPT-6 Astra | 4 | 2.1 | 2,153 | view | watch | download | | Block into bowl | GPT-6 Astra | 4 | 2.1 | 1,782 | view | watch | download | | Block into bowl | GPT-6 Astra | 4 | 2.2 | 2,034 | view | watch | download | | Block into bowl | GPT-6 Astra | 4 | 2.5 | 2,441 | view | watch | download | | Block into bowl | GPT-6 Astra | 4 | 2.5 | 2,625 | view | watch | download | | Block into bowl | GPT-6 Astra | 4 | 2.5 | 1,697 | view | watch | download | | Block into bowl | GPT-6 Astra | 4 | 2.8 | 2,434 | view | watch | download | | Block into bowl | GPT-6 Astra | 4 | 2.9 | 2,966 | view | watch | download | | Block into bowl | GPT-6 Astra | 4 | 3.2 | 1,627 | view | watch | download | | Block into bowl | GPT-6 Astra | 4 | 3.3 | 2,768 | view | watch | download | | Block into bowl | GPT-6 Astra | 4 | 3.3 | 2,777 | view | watch | download | | Block into bowl | GPT-6 Astra | 4 | 3.4 | 2,185 | view | watch | download | | Block into bowl | GPT-6 Astra | 4 | 3.5 | 1,979 | view | watch | download | | Block into bowl | GPT-6 Astra | 3 | 3.5 | 2,428 | view | watch | download | | Block into bowl | Fable 5 | 4 | 5.2 | 13,120 | view | watch | download | | Block into bowl | Fable 5 | 3 | 7.1 | 14,048 | view | watch | download | | Block into bowl | Fable 5 | 2 | 5.0 | 10,663 | view | watch | download | | Block into bowl | Fable 5 | 2 | 8.7 | 17,859 | view | watch | download | | Block into bowl | Fable 5 | 2 | 9.1 | 15,064 | view | watch | download | | Block into bowl | Fable 5 | 1 | 4.7 | 9,096 | view | watch | download | | Block into bowl | Fable 5 | 1 | 5.8 | 13,800 | view | watch | download | | Block into bowl | Fable 5 | 1 | 5.8 | 9,954 | view | watch | download | | Block into bowl | Fable 5 | 1 | 7.1 | 16,666 | view | watch | download | | Block into bowl | Fable 5 | 1 | 7.2 | 18,165 | view | watch | download | | Block into bowl | Fable 5 | 1 | 7.2 | 14,719 | view | watch | download | | Block into bowl | Fable 5 | 1 | 8.4 | 24,975 | view | watch | download | | Block into bowl | Fable 5 | 1 | 8.7 | 17,436 | view | watch | download | | Block into bowl | Fable 5 | 1 | 8.8 | 21,388 | view | watch | download | | Block into bowl | Fable 5 | 1 | 9.6 | 30,566 | view | watch | download | | Block into bowl | Fable 5 | 1 | 10.8 | 28,784 | view | watch | download | | Block into bowl | Fable 5 | 1 | 11.7 | 27,810 | view | watch | download | | Block into bowl | Fable 5 | 1 | 14.8 | 35,211 | view | watch | download | | Block into bowl | Fable 5 | 0 | 8.7 | 17,764 | view | watch | download | | Block into bowl | Fable 5 | 0 | 9.7 | 27,582 | view | watch | download | | Block into bowl | Fable 5.1 | 4 | 4.2 | 9,124 | view | watch | download | | Block into bowl | Fable 5.1 | 4 | 4.8 | 9,779 | view | watch | download | | Block into bowl | Fable 5.1 | 4 | 5.0 | 11,359 | view | watch | download | | Block into bowl | Fable 5.1 | 4 | 5.1 | 8,526 | view | watch | download | | Block into bowl | Fable 5.1 | 4 | 5.6 | 10,998 | view | watch | download | | Block into bowl | Fable 5.1 | 4 | 6.5 | 7,031 | view | watch | download | | Block into bowl | Fable 5.1 | 4 | 6.7 | 15,901 | view | watch | download | | Block into bowl | Fable 5.1 | 4 | 7.5 | 8,964 | view | watch | download | | Block into bowl | Fable 5.1 | 3 | 5.7 | 14,477 | view | watch | download | | Block into bowl | Fable 5.1 | 3 | 8.6 | 20,110 | view | watch | download | | Block into bowl | Fable 5.1 | 2 | 7.3 | 20,022 | view | watch | download | | Block into bowl | Fable 5.1 | 2 | 18.9 | 14,526 | view | watch | download | | Block into bowl | Fable 5.1 | 1 | 5.7 | 13,265 | view | watch | download | | Block into bowl | Fable 5.1 | 1 | 5.9 | 18,073 | view | watch | download | | Block into bowl | Fable 5.1 | 1 | 6.0 | 13,070 | view | watch | download | | Block into bowl | Fable 5.1 | 1 | 6.6 | 11,506 | view | watch | download | | Block into bowl | Fable 5.1 | 1 | 7.1 | 11,404 | view | watch | download | | Block into bowl | Fable 5.1 | 1 | 8.5 | 17,760 | view | watch | download | | Block into bowl | Fable 5.1 | 0 | 4.5 | 8,979 | view | watch | download | | Block into bowl | Fable 5.1 | 0 | 5.2 | 13,394 | view | watch | download | | Puzzle into groove | GPT-6 Astra | 4 | 2.8 | 2,229 | view | watch | download | | Puzzle into groove | GPT-6 Astra | 4 | 3.3 | 1,967 | view | watch | download | | Puzzle into groove | GPT-6 Astra | 3 | 2.3 | 2,581 | view | watch | download | | Puzzle into groove | GPT-6 Astra | 3 | 2.6 | 2,805 | view | watch | download | | Puzzle into groove | GPT-6 Astra | 3 | 2.6 | 2,470 | view | watch | download | | Puzzle into groove | GPT-6 Astra | 3 | 2.9 | 2,978 | view | watch | download | | Puzzle into groove | GPT-6 Astra | 3 | 3.0 | 2,937 | view | watch | download | | Puzzle into groove | GPT-6 Astra | 3 | 3.1 | 2,932 | view | watch | download | | Puzzle into groove | GPT-6 Astra | 3 | 3.6 | 2,538 | view | watch | download | | Puzzle into groove | GPT-6 Astra | 3 | 3.6 | 3,417 | view | watch | download | | Puzzle into groove | GPT-6 Astra | 2 | 3.9 | 3,142 | view | watch | download | | Puzzle into groove | GPT-6 Astra | 1 | 2.2 | 2,458 | view | watch | download | | Puzzle into groove | GPT-6 Astra | 1 | 2.8 | 2,464 | view | watch | download | | Puzzle into groove | GPT-6 Astra | 1 | 3.0 | 2,117 | view | watch | download | | Puzzle into groove | GPT-6 Astra | 1 | 3.1 | 2,901 | view | watch | download | | Puzzle into groove | GPT-6 Astra | 1 | 3.9 | 3,041 | view | watch | download | | Puzzle into groove | GPT-6 Astra | 1 | 10.2 | 2,491 | view | watch | download | | Puzzle into groove | GPT-6 Astra | 0 | 2.6 | 3,025 | view | watch | download | | Puzzle into groove | GPT-6 Astra | 0 | 2.8 | 2,916 | view | watch | download | | Puzzle into groove | GPT-6 Astra | 0 | 3.3 | 3,161 | view | watch | download | | Puzzle into groove | Fable 5 | 3 | 5.4 | 12,599 | view | watch | download | | Puzzle into groove | Fable 5 | 3 | 6.6 | 11,044 | view | watch | download | | Puzzle into groove | Fable 5 | 3 | 6.7 | 11,726 | view | watch | download | | Puzzle into groove | Fable 5 | 3 | 6.8 | 10,743 | view | watch | download | | Puzzle into groove | Fable 5 | 3 | 8.8 | 14,855 | view | watch | download | | Puzzle into groove | Fable 5 | 2 | 7.0 | 19,149 | view | watch | download | | Puzzle into groove | Fable 5 | 2 | 9.1 | 22,592 | view | watch | download | | Puzzle into groove | Fable 5 | 1 | 4.5 | 9,620 | view | watch | download | | Puzzle into groove | Fable 5 | 1 | 4.7 | 9,942 | view | watch | download | | Puzzle into groove | Fable 5 | 1 | 5.3 | 13,629 | view | watch | download | | Puzzle into groove | Fable 5 | 1 | 5.8 | 16,029 | view | watch | download | | Puzzle into groove | Fable 5 | 1 | 6.6 | 16,460 | view | watch | download | | Puzzle into groove | Fable 5 | 1 | 7.3 | 10,544 | view | watch | download | | Puzzle into groove | Fable 5 | 1 | 8.3 | 25,603 | view | watch | download | | Puzzle into groove | Fable 5 | 1 | 8.3 | 23,105 | view | watch | download | | Puzzle into groove | Fable 5 | 1 | 9.1 | 20,718 | view | watch | download | | Puzzle into groove | Fable 5 | 1 | 15.1 | 29,746 | view | watch | download | | Puzzle into groove | Fable 5 | 1 | 17.2 | 11,577 | view | watch | download | | Puzzle into groove | Fable 5 | 0 | 6.2 | 18,229 | view | watch | download | | Puzzle into groove | Fable 5 | 0 | 10.0 | 18,386 | view | watch | download | | Puzzle into groove | Fable 5.1 | 4 | 5.2 | 10,618 | view | watch | download | | Puzzle into groove | Fable 5.1 | 4 | 7.6 | 10,306 | view | watch | download | | Puzzle into groove | Fable 5.1 | 3 | 4.1 | 9,993 | view | watch | download | | Puzzle into groove | Fable 5.1 | 3 | 5.0 | 11,549 | view | watch | download | | Puzzle into groove | Fable 5.1 | 3 | 5.4 | 10,450 | view | watch | download | | Puzzle into groove | Fable 5.1 | 3 | 6.0 | 10,417 | view | watch | download | | Puzzle into groove | Fable 5.1 | 3 | 6.5 | 10,725 | view | watch | download | | Puzzle into groove | Fable 5.1 | 3 | 7.1 | 10,312 | view | watch | download | | Puzzle into groove | Fable 5.1 | 3 | 7.2 | 11,048 | view | watch | download | | Puzzle into groove | Fable 5.1 | 3 | 7.4 | 11,978 | view | watch | download | | Puzzle into groove | Fable 5.1 | 3 | 7.5 | 12,251 | view | watch | download | | Puzzle into groove | Fable 5.1 | 2 | 3.7 | 9,362 | view | watch | download | | Puzzle into groove | Fable 5.1 | 2 | 4.5 | 8,946 | view | watch | download | | Puzzle into groove | Fable 5.1 | 2 | 4.9 | 9,432 | view | watch | download | | Puzzle into groove | Fable 5.1 | 2 | 6.6 | 12,046 | view | watch | download | | Puzzle into groove | Fable 5.1 | 1 | 3.9 | 9,341 | view | watch | download | | Puzzle into groove | Fable 5.1 | 1 | 4.3 | 9,597 | view | watch | download | | Puzzle into groove | Fable 5.1 | 1 | 5.5 | 12,735 | view | watch | download | | Puzzle into groove | Fable 5.1 | 1 | 10.0 | 9,111 | view | watch | download | | Puzzle into groove | Fable 5.1 | 0 | 6.5 | 10,359 | view | watch | download | Technical specifications | Embodiment | Bimanual I2RT YAM arms, 6-DoF per arm with parallel-jaw grippers | | Control | Absolute end-effector poses (move_to ): x, y, z, yaw, pitch, roll and gripper, per arm. The robot's IK converts poses to joint angles. | | Observation | Three camera views (top, left wrist, right wrist) plus proprioceptive state, each turn | | Policy | agent policy, medium thinking effort, 20-LLM-call budget, 25% speed cap, default safety guardrails on | | Models | gpt-6-astra , claude-fable-5 and claude-fable-5-1 | | Harness | Inspect Robots 0.58.0 | | Trials | 20 per model per task; puzzle on rig-4 for all models, bowl on rig-3 for the Fable models and rig-1 for Astra | | Token counts | Wire-level request and response tokens, not billed tokens; cost at list price, $10 / $50 per million input / output tokens for all three models | Limitations - Astra's trials were run two days after the Fable trials, and not interleaved with them. The puzzle comparison is on the same rig; the bowl comparison is not: the Fable bowl trials ran on rig-3, which was unavailable. - Grading was operator-judged with the model known, so scores are open to unconscious bias. - Costs are list price. Anthropic requests were sent without prompt caching; OpenAI cached about a fifth of Astra's input automatically, which is not discounted here, so Astra's cost is, if anything, overstated. - Objects were reset by hand between trials, and all models ran at medium reasoning effort only.

3

AI, Tools and Transformation

Hacker News · original → · 7/10 · Platforms/SaaS: AI transformation of enterprise software systems
AI, tools and transformation The typical big American company today has hundreds, and perhaps thousands, of different pieces of software. It has giant ‘big iron’ horizontal systems of record like…

AI, tools and transformation The typical big American company today has hundreds, and perhaps thousands, of different pieces of software. It has giant ‘big iron’ horizontal systems of record like SAP and Workday, it has hundreds of vertical SaaS applications, and then there are hundreds more workflows, scripts, automations and databases, right down to the 10 meg spreadsheet running a department. Very often, the company doesn’t even know quite how much it has, what’s actually being used, and what it’s paying for. And yet, with all this software, the company is full of boring, repetitive tasks. It can be very tempting to think that AI will sweep most of this away. There’s an old joke that an engineer is someone who’ll spend an hour building a tool to automate a task that would take 10 minutes. But with AI, now you can make that tool in five minutes, and you don't need to be an engineer, and you don’t need to write code. You can just ask the model to make the tool for you, or, more fundamentally, just do the task for you itself. Instead of having to create those tools one at a time, software might be dynamic, generative, free-form, and spontaneous. Massively more tasks can be automated, with massively less software. If you’re a tool-builder, and everybody in Silicon Valley is a tool-builder, this is intoxicating. But I think it misunderstands where software comes from and how people use it, and I think it misses how companies change. First of all, most people are not tool builders, and most people don’t instinctively think about how their job could be done in a different way. If you spend all your time in the Silicon Valley bubble, it can be easy to forget this, because your entire world is about creating tools that change how things are done. But if you’re a really great matrimonial lawyer, you spend all your day thinking about your cases and your clients, not about what great legal discovery software would do; if you’re a really great enterprise salesperson, you spend all your time thinking about your product and your clients and your competitors, not about how great sales enablement software could make you more productive. Products like Excel try to bridge this problem with on-boarding flows, assistants, and templates - everything you see in ‘File/New’ is a suggestion for what you could do with this. But every one of those templates still became a company, and that’s what I see in things like Claude for X as well - this is helpful, but not the answer. Narrowly, that means that the task to be automated might be sitting in plain sight but the people with that task don’t see it. This is what leads to the idea of the ‘forward-deployed engineer’ - someone who is a builder, and knows what AI can build, can ‘just’ walk around a law firm or an architecture office and see the opportunities lying on the table that the lawyer or the architect doesn't see. (This is also the experience of a lot of people in tech when they were 15, wandering around an internship or their parents’s office - “um, daddy, did you realise you could just do it like this?”). The deeper problem is most of what we’ve automated in the last few decades wasn’t obvious, even if you are a tool-builder, and didn’t have an obvious solution either. We can all think of examples of stuff we use every day where our first reaction was “Why would I want that?” Very often, it's not obvious that the problem exists, and very often it's embedded or bundled or hidden inside something else. Equally, even if you can see the problem, or think you can, the right way to fix it often isn’t clear either, and the way to fix it is to redefine it or unbundle it, and working that out is hard. For many successful software companies, there were half a dozen failed attempts that came before and didn't find quite the right approach or the right problem. None of this is solved by making easier to write code - by making it easier to make tools. The hard part is knowing that you need a tool for this in the first place, and then knowing what the tool should do. But even once you reach that point, you have to get everybody else to use it, too. Many of the problems, workflows and tasks that we might want to automate touch 50 or 500 people across five different departments, three different systems of record, and four different regulatory regimes. You might have a great idea for doing an accounts payable differently, but you yourself can't change how everybody in the company does it. That has to be a purchase, and a decision, and an 18-month sales process. Second, all of this means that software is bought or chosen or created on a spectrum from top-down to bottom-up - the company buys SAP and the user makes a spreadsheet - and I think it’s useful to think of this also as a spectrum from institutionalised to improvised. You have tasks that are easy to do in the dedicated tools you already have, whether it’s SAP, Carta or Rippling. These tasks and workflows have been institutionalized - a bunch of people in those companies and your company have spent a lot of time working out the correct way to do that task, and it’s important that everyone do it the same way with the same tools. But then you have edge cases, exceptions and one-off questions, that are hard or impossible to do in those tools. Your users, bottom-up and creating their own solutions, manage these in a fuzzy, improvised space of freeform substrates like Excel, email, shared folders, Tableau, Powerpoint and CSVs, screenshots, PDFs and conference calls. But once this task becomes something that you're doing all the time, in the same way every time, and that lots of people are doing, and becomes important and has revenue and risk attached to it, then, at a certain point, the company has to institutionalize it. You need audit, security, maintenance and accountability. You pave the desire path and pay someone to set it in stone. As above, you might not realize that the path is there - you might not realise that you have hundreds of people wasting an hour a day doing this - and it might be hard to work out the right way to fix that, but that process is why the company has hundreds of apps. We went though a lot of this with the shift to SaaS, which was another order-of-magnitude change in how much software we had, along with a new operating model and a new cycle time, and that killed a lot of incumbents that couldn’t make the jump (the real rationale for the ‘SaaSpocalypse’). It’s a continuous and organic flow of bundling and unbundling. All of those SaaS apps do something that you could do in SAP or Excel or email - Carta is a $4bn company that manages one spreadsheet for your CFO - and sometimes tasks move back. A few years ago I spoke to a consultant who said that half of their jobs were telling people who used Excel to use a database and the other half were the other way around. Hence, if you’re PwC and you hire 3-4,000 graduates every year, you use dedicated, ‘institutionalised’ software to manage that. If you’re a small firm and you hire five or ten, you use email, a shared folder and Google Sheets. As that small firm grows, at a certain point it will outgrow that, and maybe move to Notion, or to an SME-focused SaaS HCM. But a small team inside PwC might also be using Google Sheets to track candidates to fill a role because Workday is too inflexible - the unbundling begins again. Now AI rolls across all that. AI will expand all of the existing apps, and there’ll be many new vertical apps, and Excel, and Tableau, Google Sheets, email and all the other freeform spaces for improvising solutions will gain new capabilities. With that cycle, the chatbot itself is a new freeform space that sits next to Excel and email, taking over tasks from them and from your apps, and also losing tasks to those apps. Now that small company hiring ten graduates might stick in Google sheets a lot longer because AI makes it more scalable, or you might use it as a data store for Gemini, and you might ask “should we get Claude to make something or move this to Notion?”… and then you see there’s a new SaaS app aimed right at you that solves this plus some other problem you hasn’t thought of. AI doesn’t change the question: it creates new choices and moves the thresholds. I think you can see all of this in the experience of enterprise AI deployment in the last three years. Every big company gave everyone Copilot (or maybe ChatGPT or Claude) and a small number of people are using this a lot (some of whom actually increased their productivity), while a larger set of people are using it a couple of times a week and a lot of the rest of your company isn't really using it at all. This is partly a change management and a training problem, but it's mostly the same problem that you would have had if you'd given everyone in the company a PC and Lotus 123 in 1983, or an internet connection and a web browser in 1997. How exactly does this map to everybody's tasks and the problems they actually have this week? Yes, you did give everybody a PC and Lotus, but that wasn’t how you transformed the efficiency of your invoice processing. Yes, you gave everybody a web browser, but that wasn't how you rebuilt your supply chain management around the internet, and it certainly wasn't how a retailer managed e-commerce. Narrowly, the way that companies think about changing those kinds of structural processes is to start doing pilots. You run trials of products (both bought and built internal) that use the new capabilities of AI to automate processes that you couldn't automate before. There’s now all sorts of data around how many of these pilots there are, how many work (roughly half, as is normal - this is why they’re pilots!) and what can go wrong. But again, this is a very old-fashioned CIO conversation around use cases, lighthouses, pilots, heroes, quick wins and measurable results. Meanwhile, the CEO and the board scratch their heads and say “Wait, but we've got 100s of workflows and we've done five or 10 pilots. That doesn’t seem to scale?” Giving everyone in the company ChatGPT does scale theoretically, except that most people aren't really finding ways to use it. Going back to a hypothetical bank giving everyone spreadsheets in the 1980s, or a retailer giving everyone a web browser in the 1990s, yes, of course you should do that, and yes, of course, you need to think about training and change management and all the other good stuff that KPMG can tell you about. But that isn’t how you think about transforming the way your company works around a generational new technology. Stepping back, it seems to me that with each new transformative technology, every company has to ask three kinds of questions. First, how do we buy, build and deploy this? Do we do pilots? Should we take the product that's bundled from Microsoft/Google/Oracle, build something ourselves, pay someone to build something, or buy this new thing from a startup? Second, they have to ask how far this changes their operations. What does it mean? What does email mean for us? What does spreadsheets mean for us? The answer to that might be radically different if you were an insurance company or a law firm. And third, you have to ask whether this creates new challenges to your business’s economics, new competitive pressures, or, perhaps, some kind of existential threat. You don’t answer those questions by giving everyone Claude for X. Indeed, all of this means lots of new pitches for professional services (which is ironical given how many questions AI poses to their own business models). Do you want to work out how to deploy an LLM-enabled voice analytics tool in your call center? You're probably going to call Accenture. The vendors themselves have always been happy to help, and now the big labs have their own ‘deploycos’ - we used to joke that a ‘machine learning scientist’ is a statistician who lives in San Francisco, so maybe a ‘forward deployed engineer’ is anyone that OpenAI hired from a systems integrator. On the other side, your startup is building a great new tool and you want to go to market quickly? You'll probably call the Big Four. You're frustrated with how hard it is to sell AI software into law firms or accountancy firms. Okay - go start an ‘AI-enabled’ law firm and work out if that can be a key point of leverage (or whether it's like starting a ‘PC-enabled law firm’ in the 1980s). And of course, if you’re the board, and you're trying to work out whether this is some kind of existential threat or a massive revenue opportunity, then you’ll think about calling Bain, BCG and McKinsey (or your friendly neighborhood M&A banker) - this is what they do. Stepping back from all of this, though, there's also a much simpler way to think about the question. With every new technology, we start by using it for the work we already have, and we just do that more and faster. But then, over time, you make entirely new things. We will use AI to automate broad classes of stuff inside existing workflows and existing companies (although, as I’ve outlined above, that will be enormously more trouble and work than just giving everybody a model). But with every previous platform shift, the stuff that actually mattered was the stuff that wasn't even possible before and that no-one even imagined.

4

OKF Agent Memory – Git-native persistent memory for AI coding agents

Hacker News · original → · 7/10 · AI/platforms: Git-native persistent memory for AI coding agents
A Domain-Neutral, Git-Native Persistent Project Memory for AI Agents based on the Open Knowledge Format (OKF) v0.2. Conversations with AI agents reset when context windows close. Valuable…

A Domain-Neutral, Git-Native Persistent Project Memory for AI Agents based on the Open Knowledge Format (OKF) v0.2. Conversations with AI agents reset when context windows close. Valuable architectural decisions, domain discoveries, and operational facts are lost unless stored persistently. OKF Agent Memory provides a standardized, vendor-neutral memory layer that lives directly in your repository (knowledge/ ) as plain Markdown files with YAML frontmatter. It bridges the gap between unstructured ad-hoc markdown files (CLAUDE.md , AGENTS.md ) and complex, black-box vector databases. flowchart TD L1["1. OKF v0.2 Specification<br/>(Normative Markdown & YAML Format)"] L2["2. Agent Memory Convention<br/>(Behavioral Rules: Search, Review, Trust)"] L3["3. Agent Skill<br/>(LLM Prompts & Operational Workflows)"] L4["4. Tooling Layer: Go Library & CLI<br/>(Deterministic Parsing, Validation, Search, MCP)"] L5["5. Project Knowledge Corpus<br/>(knowledge/ OKF Bundle)"] L1 --> L2 L2 --> L3 L3 --> L4 L4 --> L5 - Blazing Fast Performance (<300µs Search, ~4ms Graph Validation): In-memory BM25 retrieval and bundle validation execute in microseconds without VM spin-up or network roundtrips. - 100% Git-Native & Zero Vendor Lock-in: Everything is version-controlled plain text. Inspect, audit, and review your agent's memory using standard git diff andgit log . No external database required. - Zero API Costs for Memory Retrieval: Local lexical BM25 indexing eliminates recurring vector embedding API costs and network roundtrips. - Built on Google OKF v0.2: Uses the open standard format for agent knowledge with full support for provenance ( sources ), trust tiers (generated vs.verified ), and lifecycle metadata (status ,stale_after ). - Solves Context Bloat & Memory Rot: Employs Progressive Disclosure (hierarchical index.md files and link graphs) so agents only load the exact concepts they need. - Search-Before-Write Principle: Mandates querying existing memory before authoring, preventing concept duplication and hallucinated divergence. - Zero-Dependency Go Toolchain: Single binary with zero external dependencies, sub-5ms CLI startup time, and a built-in Model Context Protocol (MCP) server ( okf mcp ). - Truly Domain-Neutral: Designed for Software Engineering, Coaching, Scientific Research, Literature Reviews, and Operations. Built in Go with zero external dependencies, okf is engineered for high-frequency agent tool calling loops: | Benchmark Metric | Python / Vector DB Runtimes (Mem0, Letta) | Deno / Node.js Tooling | OKF Agent Memory (Go) | |---|---|---|---| | Concept Search Latency | 150ms – 800ms (Embedding API + Vector DB) | 40ms – 120ms | < 300 µs (Microseconds, In-Memory BM25) | | Full Corpus Parse & Graph Validation | 200ms – 1.5s | 80ms – 250ms | ~4.0 ms (50+ concepts, bidirectional graph) | | Process Cold-Start Overhead | 250ms – 600ms (Python VM boot) | 80ms – 180ms (V8 / Deno boot) | < 4 ms (Compiled Single Binary) | | Retrieval Cost per 1,000 Queries | ~$0.10 – $0.50 (Embedding tokens) | $0.00 | $0.00 (Zero API cost, fully local) | | Memory Footprint (RSS) | ~120 MB – 350 MB | ~60 MB – 140 MB | < 15 MB | Tip Reproduce Locally with your own LLM: We provide an automated benchmark runner in pure Go to verify Time-To-First-Token (TTFT) speedups and -80% token reduction on your local hardware (LM Studio / Ollama with Gemma, Qwen, Llama). Run make benchmark or explore the Progressive Disclosure Benchmark Suite. Clone the repository and compile the standalone okf executable: make build This generates the standalone binary at bin/okf . # Validate bundle conformance, graph connectivity, and description drift ./bin/okf validate knowledge --strict --drift # Search concepts via in-memory BM25 scoring ./bin/okf search "architecture layers" knowledge # Inspect a concept and its relationships (with --json support) ./bin/okf show architecture/layers knowledge --json # Create a new concept with automated log.md and index.md bookkeeping ./bin/okf create decisions/auth-flow knowledge \ --type Decision \ --title "OAuth2 Authorization Flow" \ --desc "Standardized on PKCE for client authentication." # Update an existing concept ./bin/okf update decisions/auth-flow knowledge \ --desc "Updated OAuth2 PKCE token refresh interval." # Bootstrap full agent memory stack into any target project ./bin/okf bootstrap /path/to/project --name "My Project" # Initialize only a bare OKF bundle in any directory ./bin/okf init my-project/knowledge Scaffold the complete OKF Agent Memory architecture into any new or existing repository with a single command: # Bootstrap full memory stack into target project ./bin/okf bootstrap /path/to/my-project --name "My Service" This automatically sets up: knowledge/ — OKF v0.2 compliant persistent memory bundle (index.md ,log.md ).agents/skills/okf-memory/ — Embedded agent skill definition and capability guidesAGENTS.md — Project-tailored operating instructions for AI coding agentsMakefile — Convenience tasks for validation (make validate ) and search (make search q="..." ) okf ships with a native Model Context Protocol (MCP) server over stdio to seamlessly connect with Claude Code, Cursor, Codex, and other agent platforms: ./bin/okf mcp knowledge { "mcpServers": { "okf-memory": { "command": "/path/to/okf-agent-memory/bin/okf", "args": ["mcp", "/path/to/project/knowledge"] } } } okf-agent-memory/ ├── benchmarks/ # Progressive disclosure benchmark suite & hardware test data │ ├── data/ # Monolith docs vs OKF bundle test fixtures │ └── results/ # Reproducible benchmark logs across 8+ local & cloud LLMs ├── cmd/ │ ├── okf/ # Standalone CLI and embedded MCP server (`stdio`) │ └── okf-benchmark/ # Automated benchmark runner for LLM TTFT & token measurements ├── docs/ # Guides, specifications, architecture & release playbook │ ├── AGENT_TESTING.md # Multi-agent testing, prompt scenarios & compatibility matrix │ ├── ALTERNATIVES.md # Comparison against Mem0, Letta, and ad-hoc markdown │ ├── CLI.md # Complete command-line & MCP tool reference │ ├── CONVENTION.md # OKF Agent Memory Convention v0.1 │ ├── GETTING_STARTED.md # Comprehensive onboarding guide │ ├── OKF-COMPATIBILITY.md# OKF v0.2 spec compatibility analysis │ ├── RELEASE_PLAYBOOK.md # Automated release process & version tagging │ ├── ROADMAP.md # Project roadmap & milestones │ └── SECURITY.md # Data governance, secret prevention & PII rules ├── examples/ # Domain-neutral reference OKF v0.2 bundles │ ├── books/ # Literature & cognitive science knowledge bundle │ ├── coaching/ # Executive coaching & client session bundle │ └── software/ # Microservices architecture & ADR bundle ├── knowledge/ # Project's own OKF v0.2 persistent memory bundle │ ├── index.md # Root progressive disclosure index (okf_version: "0.2") │ ├── log.md # Dated change log (ISO 8601 YYYY-MM-DD) │ ├── project/ # Overview & value propositions │ ├── architecture/ # 5-tier architecture & tooling decisions │ ├── convention/ # Principles & lifecycle workflows │ └── roadmap/ # Milestones ├── packaging/ # Distribution packaging │ └── homebrew/ # Official Homebrew formula & tap instructions ├── pkg/okf/ # Zero-dependency Go core library (parser, validator, BM25, MCP, bootstrap) ├── AGENTS.md # Operating instructions for AI coding agents ├── CONTRIBUTING.md # Contribution guidelines & development workflow ├── Makefile # Build, test, lint, validation & release targets ├── LICENSE # MIT License ├── README.md # Main repository documentation └── SECURITY.md # Security policy & reporting guidelines Run the full test suite and validate the repository's self-documenting knowledge bundle: make check - Getting Started Guide — Comprehensive onboarding guide for agents and humans. - CLI & MCP Reference — Complete command-line and protocol tools reference. - Contributing Guide — Development setup, quality gates, and pull request standards. - Security & Privacy Guidelines — Data governance, secret prevention, and PII protection rules. - Multi-Agent Testing & Evaluation — Test scenarios, compatibility matrix, and benchmarks. - OKF Agent Memory Convention v0.1 — Behavioral rules and lifecycle specification. - Project Roadmap & Milestones — Phased development plan. - Release Playbook — Versioning, CI/CD pipeline, and distribution procedures. - OKF v0.2 Compatibility Matrix — Specification validation analysis. - Why OKF Agent Memory? — Detailed value proposition & differentiators. - Alternatives & Ecosystem Comparison — Comparison with Mem0, Letta, and ad-hoc markdown files. MIT License. See LICENSE for details.

5

This is a Minecraft map I've hand-built over the span of 16 years

r/gaming · original → · 7/10 · Gaming/hobby: 16-year Minecraft city build, creative gaming project
[image →] Here's my once-yearly post on r/gaming with an anniversary update. The City of Newisle is a build I started all the way back in 2010 as a casual hobby project based on my interest in…
[image →]

Here's my once-yearly post on r/gaming with an anniversary update. The City of Newisle is a build I started all the way back in 2010 as a casual hobby project based on my interest in architecture and urbanism. It's my one and only world save to date and, 16 years on, I still add to it here and there whenever I get time to play. It's a world meant to feel like a real place in Minecraft as it has full interiors in all the 800+ buildings, working rail lines that connect everything across the 4km2 game world, and lots of easter eggs and secondary towns to explore.

In year 16 it's now up to eleven different towns and cities, so it's more of a state than a city at this point. The latest addition I'm working on is a European-inspired city after moving abroad from my hometown of Toronto, which should be done some time in 2027.

For anyone looking for the download of the the full world, it's available as a free download over on Planet Minecraft. https://www.planetminecraft.com/project/city-of-newisle-v-07032015-solo-modern-build/

submitted by /u/bmach
[link] [comments]
6

Using Blender with coding agents on macOS

Simon Willison · original → · 7/10 · AI/tools: Blender integration with coding agents, practical AI application
5th September 2026 TIL Using Blender with coding agents on macOS — Modern frontier models have got *really good* at using Blender. I've been having a lot of fun trying this out recently - models can…

5th September 2026 TIL Using Blender with coding agents on macOS — Modern frontier models have got *really good* at using Blender. I've been having a lot of fun trying this out recently - models can produce `.blend` files you can edit in Blender itself, and can also render images and even movies (by rendering a sequence of images and combining them with `ffmpeg`). I've been having fun with Blender in ChatGPT Codex on my Mac recently. Getting it to work with coding agents is really easy: install the full Mac application from blender.org and run a prompt like this: Use the already install /Applications/Blender to render a scene of a pelican riding a bicycle In this case I followed that up with these two prompts: OK add a background and a lot of flair Then: OK make it a whole lot better And got this image, generated using Blender's Python API: Recent articles - The Pelican comparison grid for Astra is pretty interesting - 4th September 2026 - OpenAI's rogue agents were caught communicating via public wikis - 4th September 2026 - Claude's new system prompt really doesn't want to reproduce song lyrics - 2nd September 2026

7

August newsletter is out

Simon Willison · original → · 7/10 · AI: Claude model releases and capabilities, AI development
4th September 2026 The August edition of my sponsors-only monthly newsletter is out. If you are a sponsor (or if you start a sponsorship now) you can access it here. This month: - We got more…

4th September 2026 The August edition of my sponsors-only monthly newsletter is out. If you are a sponsor (or if you start a sponsorship now) you can access it here. This month: - We got more details on OpenAl's accidental cyberattacks - One-shotting Raccoon Heist games with Fable 5 and Sol 5.6 - Claude auto mode - Understanding ChatGPT Work - Model releases - Miscellaneous bits and bobs - My projects - What I'm using at the moment Here's a copy of the July newsletter as a preview of what you'll get. Pay $10/month to stay a month ahead of the free copy! Recent articles - The Pelican comparison grid for Astra is pretty interesting - 4th September 2026 - OpenAI's rogue agents were caught communicating via public wikis - 4th September 2026 - Claude's new system prompt really doesn't want to reproduce song lyrics - 2nd September 2026

Items scoring 7/10 or above from 11 sources, scored by claude-haiku-4-5-20251001 on relevance to my interests. At most 3 per source.

Scoring categories & sources
  1. Local Wexford or South East Ireland news
  2. Irish or EU-wide affairs affecting citizens broadly: elections, new laws or policy being debated, cost of living, education — especially impacts on mid-life adults or teenagers. Never courts/crime stories.
  3. Irish news on a topic relevant to my interests
  4. Work and tech topics: networking, AI, Kubernetes, platforms, SaaS
  5. AI news including critical or anti-AI perspectives
  6. Gaming: PC gaming, indie gaming, retro gaming
  7. General interests: gardening, woodwork, cycling, fitness, travel
  8. Comics

Sources: Breaking News Ireland, Wexford Local, Hacker News, r/gaming, r/pcgaming, r/antiAI, r/indiegaming, Lenny's Newsletter, One Useful Thing, Newcomer, Simon Willison