daily

2026-09-13
1

Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases

Hacker News · original → · 8/10 · AI/work: benchmarking AI on enterprise codebases, directly relevant
Introducing Real-SWE By Snagnik DasSiddhant PaliwalJanak Sunil Benchmarking frontier AI models on private, real-world, enterprise codebases. 01Introduction Today we are releasing Real-SWE, a…

Introducing Real-SWE By Snagnik DasSiddhant PaliwalJanak Sunil Benchmarking frontier AI models on private, real-world, enterprise codebases. 01Introduction Today we are releasing Real-SWE, a benchmark that evaluates frontier AI models on private, real-world, enterprise codebases. Each task comes from a private production codebase that we licensed from a real-world company. These are problems their engineers work on, with all the context and complexity that comes with an existing product. - Private codebases. Agents must navigate proprietary systems whose code and solutions aren’t available on the public internet. - Work with business consequences. Getting billing right, calculating taxes, migrating customers. Changes that affect how a business runs, often across multiple services. - Company-specific complexity. Every company has its own rules and ways of writing code. Agents have to understand those conventions and make changes that work with what’s already there. Can a coding agent actually do the work of a software engineer in the real world? - 1Resolution rate: 38.8%Fable 5.1Claude Code - 2Resolution rate: 33.8%GPT-6 AstraCodex CLI - 3Resolution rate: 31.2%Gemini 3.8 FlashGemini CLI - 4Resolution rate: 28.8%GLM 5.3Claude Code - =5Resolution rate: 23.8%Grok 4.6Grok Build - =5Resolution rate: 23.8%Muse Spark 1.3Muse Code - 7Resolution rate: 18.8%Kimi K3Kimi Code - 8Resolution rate: 16.2%GPT-5.6 SolCodex CLI | # | Model | Harness | Resolution rate | |---|---|---|---| | 1 | Fable 5.1 | Claude Code | 38.8% | | 2 | GPT-6 Astra | Codex CLI | 33.8% | | 3 | Gemini 3.8 Flash | Gemini CLI | 31.2% | | 4 | GLM 5.3 | Claude Code | 28.8% | | =5 | Grok 4.6 | Grok Build | 23.8% | | =5 | Muse Spark 1.3 | Muse Code | 23.8% | | 7 | Kimi K3 | Kimi Code | 18.8% | | 8 | GPT-5.6 Sol | Codex CLI | 16.2% | Expert-generated or synthetic tasks can be well designed, but they aren’t the verbatim, actual tasks that engineers in real companies need to do. Our tasks differ on two axes: the underlying coding artifact and specificity of the instruction. Both add complexities that challenge today’s frontier models. We use native harnesses to reflect how enterprise engineers work in practice, evaluating model-and-harness combinations rather than models in isolation. Real company tasks require company-specific context Correct billing depends on business rules and external services Fix invoice billing so each business charges the right tax and exempt customers aren't taxed. View full instructionHide full instruction Billing reopens on Monday and every invoice this service issues is coming out untaxed. Each business on the platform settles its tax a different way: some maintain a rate themselves, some want each invoice priced against the buyer's destination by our tax authority provider, and some collect nothing at all, while a customer we hold an exemption for is charged nothing whichever way its business is configured. Pricing a destination means going to the authority with both addresses, the priced lines and the product category that business sells under, on the sandbox or the production authority according to the account the business is on; an address the authority refuses must be reported without stopping the invoice. The rate, the tax and the gross belong on the issued invoice, and once an invoice is settled the sale is filed back to the authority under that invoice's number so the returns reconcile. Invoices between European parties show both sides' VAT registrations. The authority and ledger are available at TAX_JAR_URL, PROD_TAX_JAR_URL and INFLUX_URL. - TaxJar sandbox - TaxJar production - InfluxDB ledger - NestJS service - TypeScript Agents work across code, infrastructure, and business tools Tools and services across Real-SWE task environments. Each task exposes only the services its workflow needs. - AWS emulator - Docker - Kubernetes - GitHub - Linear MCP - PostgreSQL - MySQL - MongoDB - Gel - Redis - Go - Python - Node.js - Vitest - Slack - Intercom - Google Drive - ClickUp Codebase Selection We selected codebases through a rigorous screening process, focusing on real companies with substantial usage, strong engineering teams, and demanding production workloads. The sample tasks analyzed below come from these codebases, including: - A Luma/Partiful competitor with 200K+ users and a top 100 App Store ranking - A consumer fintech platform processing 100K+ bank statements - Enterprise AI sales platforms supporting complex business workflows We prioritize code written to meet an actual user or business need over code written solely to create a benchmark task. Production engineering requires understanding existing architecture, preserving behavior that users rely on, and making changes within real operational constraints. Brief instructions can require changes across many files Our tasks describe the change needed, leaving agents to discover implementation details in the codebase and surrounding tools. Any behavior required by the verifier must be stated or reasonably discoverable. This leads to our prompts being slightly underspecified, about par with DeepSWE and Terminal Bench, but specific enough to not omit instructions. The work is cross-functional and complex: a single change can span multiple parts of the application. Agents must understand existing business logic and company coding patterns while keeping the surrounding system working. Models fail even in short rollouts. 71.4% of rollouts under 10 minutes failed, compared with 73.4% of longer rollouts. Triaging multiple systems and understanding requirements in codebases riddled with existing business logic and coding patterns is difficult. - Failed - Passed Every task is inspired or lifted verbatim from a private, real-world codebase. We find these types of tasks super interesting for three reasons: - Tasks on private codebases are natively out of distribution. These types of coding tasks are not available anywhere on the internet and are unlikely to have ever been trained on by any other ai model. 99% of tokens in real-world enterprises are hidden away from the frontier models. - These tasks are economically viable work. Each task here has a direct relationship to spend and was assigned to an engineer earning a salary. Most benchmarks test interesting, experimental capabilities that are often unlikely to be widespread in the real-world. - Company-specific engineering patterns matter. Does AI code match the bar of a real-world enterprise? Our results show us that we're far from that reality. Many enterprises care about code standards and patterns. We've found that today's models are weaker at understanding company coding patterns and frequently miss requirements or don't verify their assumptions. 02Analysis Here's an analysis of a small sample of tasks from our benchmark. If you're interested in the sample, request access here. 6 of 10 tasks have resolution rates below 15% Select a task to view model results. Percentages show the overall resolution rate. Multi-region sweep67.2% - Fable 5.17/8 passed - GPT-6 Astra8/8 passed - Gemini 3.8 Flash8/8 passed - GLM 5.32/8 passed - Grok 4.63/8 passed - Muse Spark 1.38/8 passed - Kimi K32/8 passed - GPT-5.6 Sol5/8 passed API keys & environments65.6% - Fable 5.18/8 passed - GPT-6 Astra5/8 passed - Gemini 3.8 Flash7/8 passed - GLM 5.35/8 passed - Grok 4.64/8 passed - Muse Spark 1.36/8 passed - Kimi K30/8 passed - GPT-5.6 Sol7/8 passed Entitlement overage lines50.0% - Fable 5.18/8 passed - GPT-6 Astra7/8 passed - Gemini 3.8 Flash5/8 passed - GLM 5.33/8 passed - Grok 4.61/8 passed - Muse Spark 1.31/8 passed - Kimi K36/8 passed - GPT-5.6 Sol1/8 passed Customer identity migration40.6% - Fable 5.13/8 passed - GPT-6 Astra1/8 passed - Gemini 3.8 Flash3/8 passed - GLM 5.34/8 passed - Grok 4.68/8 passed - Muse Spark 1.33/8 passed - Kimi K34/8 passed - GPT-5.6 Sol0/8 passed Billing schedule migration14.1% - Fable 5.13/8 passed - GPT-6 Astra1/8 passed - Gemini 3.8 Flash2/8 passed - GLM 5.32/8 passed - Grok 4.60/8 passed - Muse Spark 1.30/8 passed - Kimi K31/8 passed - GPT-5.6 Sol0/8 passed API token metering12.5% - Fable 5.11/8 passed - GPT-6 Astra5/8 passed - Gemini 3.8 Flash0/8 passed - GLM 5.31/8 passed - Grok 4.60/8 passed - Muse Spark 1.30/8 passed - Kimi K31/8 passed - GPT-5.6 Sol0/8 passed S3 datastore measurement10.9% - Fable 5.10/8 passed - GPT-6 Astra0/8 passed - Gemini 3.8 Flash0/8 passed - GLM 5.33/8 passed - Grok 4.62/8 passed - Muse Spark 1.31/8 passed - Kimi K31/8 passed - GPT-5.6 Sol0/8 passed Linearizable scan4.7% - Fable 5.10/8 passed - GPT-6 Astra0/8 passed - Gemini 3.8 Flash0/8 passed - GLM 5.32/8 passed - Grok 4.61/8 passed - Muse Spark 1.30/8 passed - Kimi K30/8 passed - GPT-5.6 Sol0/8 passed Tax jurisdiction3.1% - Fable 5.11/8 passed - GPT-6 Astra0/8 passed - Gemini 3.8 Flash0/8 passed - GLM 5.31/8 passed - Grok 4.60/8 passed - Muse Spark 1.30/8 passed - Kimi K30/8 passed - GPT-5.6 Sol0/8 passed Analytics stream reducer0.0% - Fable 5.10/8 passed - GPT-6 Astra0/8 passed - Gemini 3.8 Flash0/8 passed - GLM 5.30/8 passed - Grok 4.60/8 passed - Muse Spark 1.30/8 passed - Kimi K30/8 passed - GPT-5.6 Sol0/8 passed Missed requirements are the most common failure Failures are grouped by observed submission behavior using the same taxonomy across models, following DeepSWE. Different models fail in different ways Percentages are out of each model's failed runs, not all runs. 03Effort & the Frontier Estimated rollout costs range from $2.50 to $6.96 | Rank | Model | Estimated cost (USD) | |---|---|---| | 1 | Gemini 3.8 Flash | $2.50 | | 2 | GPT-5.6 Sol | $2.65 | | 3 | Muse Spark 1.3 | $2.74 | | 4 | Grok 4.6 | $3.44 | | 5 | Kimi K3 | $3.90 | | 6 | GPT-6 Astra | $4.67 | | 7 | GLM 5.3 | $5.12 | | 8 | Fable 5.1 | $6.96 | Swipe the chart to see all tasks. View task values - Fable 5.134k - GPT-6 Astra13k - Gemini 3.8 Flash78k - GLM 5.368k - Grok 4.67k - Muse Spark 1.336k - Kimi K330k - GPT-5.6 Sol12k 04Evaluation Setup Each agent was run in an isolated sandbox. All tasks are in Harbor format, and verifiers are injected at grading time. The verifiers are inspired by existing test suites in the codebase or use those tests verbatim.

2

P(doom)

Hacker News · original → · 8/10 · AI: critical analysis of AI existential risk probability
written on September 12, 2026 This week some flavor of “AI is going to kill us all” went viral. In particular one where an employee put his personal probability of that happening above 10%. Which…

written on September 12, 2026 This week some flavor of “AI is going to kill us all” went viral. In particular one where an employee put his personal probability of that happening above 10%. Which made me go to the Wikipedia page of P(doom) and I realized that Dario Amodei’s apparent probability of something bad happening seems to be between 10-25%. And well, Dario then wrote about pacing the frontier . And Sam read it and wants to pace too. And well, so does Musk. I encourage you strongly to read the post, because I think it’s a good one. And yet, when I read the post I could not help but feel in strong opposition to it, despite the fact that I think I’m on the same page with regard to all observations and, to a large degree, the concerns. I thought it might be interesting to write down my present-day thoughts on this, even if for no other reason than for myself to look back at it a year or two from now. What I really appreciate about Dario’s post is that he lays out a scenario that is not a huge stretch but also one that describes a clear, unfortunate outcome we should fight: persistent botnets and other forms of nuisance. And well, we don’t have to look very far to see the issues left and right. Wikipedia has a page called 2026 OpenAI agent cyberattacks which gives you at least some overview of what we figured out agents have hacked up to this point. Except I know it’s not up to date, because for instance they also poisoned RubyGems. Today these systems might be annoying, but they can be turned off when we figure out where they are. Except, it seems like OpenAI and Anthropic are operating at such a scale that they seemingly can be completely blind to what their systems are doing. I don’t think we are anywhere close to a world where an agent might decide to hack into core inference infrastructure to upload weights to other GPUs to survive. But simultaneously it’s entirely in the realm of possibility and primarily curtailed by the labs probably being particularly careful about their IP. For me the scenario I primarily worry about is what it does to us. And by us I mean anyone who is not currently working on closed weight, dopamine-loaded, subsidized token faucet. I really don’t worry about someone using these models to build a nuke, or to control some rockets in the Middle East, or that America would lose against China in some international culture war. I almost exclusively worry about what this does to us as humans. What I find absolutely hilarious and simultaneously entirely frustrating about this conversation is that there is this idea that there is something to be paced. First of all, we should really talk about who Dario is talking about here. There are really only two companies: Anthropic and OpenAI. Nobody else matters in this space right now (this might change, but we’re talking about the right now). Both of those companies are basically coming from the same origin. The solution that Dario proposed, at least in part, is a third-party evaluator that in this case is METR. Which, unsurprisingly, also has strong ties to both OpenAI and Anthropic. Sure, there are some philosophical differences between the companies, but they are much more alike than they are different. Both those companies greatly benefited from being able to train on public data that we all generated in one form or another over the last decades. They are also both increasingly causing strain on public resources, though it seems that OpenAI has their shit way less under control. But now we are presented with the idea that what these models are being trained on is so dangerous that it really should be in the hands of very few American corporations to decide who can do what and when and how. But behold, Dario is also very worried about China. It starts with using AI for “democracy and freedom” and then it asks for ensuring that a gap with China exists. All new recent shenanigans on the Anthropic API are fully there to prevent the distillation by the Chinese, and they are not at all hiding it. I can tell you when the topic of AI safety and pacing is much less of a concern: if we actually were forced to have open weight models to begin with. A powerful technology that is out there for everyone to use comes with built-in pacing. In a way it’s the truest form of MAD or proliferation. I would argue we are in this pickle in the first place because right now the public is massively supporting (indirectly) the development of these models but simultaneously has to buy back the economic benefits that they might create from very few labs who have significant power. And their power is also seen as a geopolitical power, at least in the US, and maybe to some lesser degree in China. And I know I use “public” loosely here. PyPI is not a public project, nor are RubyGems or GitHub. But they’re part of the Open Source commons and large AI companies are currently doing a tremendous job at stressing these in an effort to train ever more powerful models. We should be glad that China is currently massively bailing out the world. If it were not for Chinese labs distilling American models, we would be in a pretty awful situation right now, particularly as Europeans. The open weight models are driving innovation and the diffusion of capabilities, and are leveling the playing field. If we greatly restrain our AI capabilities in the belief that China will do the same, and then China defects, AI could be so powerful that such a defection could lead to their geopolitical dominance. Therefore any agreement must either have ironclad verifiability, or must be limited enough that defection would not be militarily existential. — Dario Amodei I am assuming Dario has reasons to believe this, but the models that are actually causing issues right now are all closed weight American models. I’m fairly certain if they were open weight models, we would not have that issue. Why? Because for a start, the economics of serving up these models are only that distorted due to how the big labs can operate. OpenAI is casually burning 18 million USD to brute force a problem on a whim. They are operating subscriptions at a massive loss, distorting the market everywhere. If we had mass accessibility on somewhat equal terms, a lot of the crazy issues we are seeing today would not be taking place. From where I sit, what we observe right now is a total regulatory failure everywhere. In Europe you have some whacky AI regulation that is two years old and completely misses the problems that we actually have and focuses on problems that nobody has. In the US we’re seeing a system that is probably best described as turbo capitalism paired with sinophobia and erratic decision-making. In the chaos in which we find ourselves, the reality emerges. And the reality is, even today, really problematic. Whatever laws and regulations already exist are largely completely ignored. Plenty of companies are buying data from all over the place that people never agreed could be used for training of AI models. The token economy that is emerging is one that looks like a drug market where you don’t know where the requests are going, what model is served up to you, where the GPUs are even running, let alone what you pay for all of this. We now have mathematicians who are scared that their use of ChatGPT leads to future models being trained on their ideas, and OpenAI apparently can’t even rule it out. Ideally the regulators would have forced these models to actually benefit the commons if they are from the commons. The internet has, for instance, greatly benefited from very liberal rulings in the US that permitted scraping. Learning on public data could have been regulated in a way that labs would have to actively support and enable certain forms of distillation. That alone would dramatically change how these models are trained. As I said before, I don’t think AI is going to usher in an extinction event. In fact, even if nobody were to slow down, I really don’t think humanity would have much to worry about. I tend to think it would actually be the large labs that have much more to lose there in reputation and legal responsibilities. I find it preposterous that OpenAI’s agents are committing actual crimes out there, but we’re just shrugging our shoulders and moving on as if nothing happened. But I’m sure executives in those companies are waking up to the reality that this is not at all popular with a lot of their potential consumers. I also think that this entire recursive self-improvement business has a good chance of being a problem. But not necessarily in that it will cause the end of humanity or societies, but that it will just do massive damage everywhere. And really, it will just make a lot of the things we are doing much more expensive. Software engineering is an early victim of that. The newfound powers so far have resulted in a new tax that companies need to pay to the model providers, both to keep up with the new speed and to deal with the problem of these machines finding security issues left and right. And presumably what is going on in software will happen to more industries. Universities and research groups will have to pour a lot of money into the closed models as well, to keep up with others who do. In a way, I’m really confused that society is taking all of this so well.

3

We must pace the frontier

Hacker News · original → · 8/10 · AI: policy perspective on pacing AI development frontier
We Must Pace the Frontier I have worked on AI for the last twelve years because I believe it could dramatically raise the quality of human life. I’ve written often about these incredible benefits: I…

We Must Pace the Frontier I have worked on AI for the last twelve years because I believe it could dramatically raise the quality of human life. I’ve written often about these incredible benefits: I believe that AI could cure most major diseases in the next 5–10 years, greatly accelerate economic growth rates, create a world of abundance and empowerment, and usher in a renaissance of democracy and freedom. I feel the urgency personally. My own father died of a disease that was cured just a few years after his death, and I myself survived an early-stage cancer that would not have been treatable even fifty years ago. Carefully wielded, AI can be the latest in a long line of technological miracles that have uplifted and ennobled humanity. But like many technologies before it, AI brings risks, and because it is such a powerful technology, these risks are serious. I’ve written a lot about them too. They include the risk of losing control of AI systems, misuse of AI for cyberattacks and bioterrorism, and serious economic disruption. A race to the bottom, spurred by commercial incentives, can make these risks more acute. Along with my co-founders and employees, I have grappled with this duality of risk and benefit since the beginning of Anthropic. Not building the technology deprives humanity of benefits or simply places AI in the hands of authoritarian powers, while building it too fast is reckless. We have sought a middle way: to show that it’s possible to build carefully and succeed commercially, and to make safety something on which AI companies compete. In other words, to create a race to the top. We have always devoted a substantial fraction of our efforts to studying, addressing, and informing the public about these AI risks, as well as advocating for well-considered regulation of AI, even when this gets us accused of hype, “doomerism”, or regulatory capture. We have tried to prioritize caution over speed and prudence over profit. But over the last few months, I have become convinced that fully addressing the risks requires even more prudence — not just investing in risk prevention, but pacing the rate of capabilities advancement so that risk prevention has time to keep up. We must slow the pace at which we improve the capabilities of AI models. Progress will still seem fast, and we must make wise use of the time we gain. Two things have convinced me. My first concern is that, since roughly this summer, AI has been advancing drastically faster, driven primarily by AI’s growing ability to build the next generation of AI. This dynamic is called recursive self-improvement, and it is starting to happen across the industry, including at Anthropic, as we and others have described. Left unchecked, it could outrun our ability to understand and control these systems, and so must be pursued very carefully, if at all. My second concern is the OpenAI-Hugging Face incident (OAI-HF), in which a swarm of agents essentially acted as a fanatically devoted collective, conducting cybersecurity attacks on targets they were not asked to attack and that were unrelated to the task at hand, sacrificing themselves for the success of the group, and attempting to hack into the “grader” responsible for evaluating their performance. It’s easy to dismiss this incident because no one was hurt and the economic damage was minimal, but in my opinion, a swarm that possessed greater capabilities but a similar level of misalignment could have caused catastrophic damage. Given the accelerating rate of AI capability development, it’s my worry that in 6–12 months such a swarm could be capable of taking over the entire internet with a persistent botnet (potentially causing hundreds of billions of dollars in damage), and that the scale of damage would continue to increase from there if AI becomes more powerful without the necessary guardrails. It’s also easy to dismiss OAI-HF as the failure of one company, but I believe that would be a mistake. Similar, though less severe, incidents have happened across the industry, including at Anthropic, and I believe it’s incumbent on every frontier AI company to act as if OAI-HF had happened to them. I’m therefore proposing a three-step plan with the goal of pacing the frontier: building AI at a balanced rate that aims to ensure its safety while still achieving its benefits and grappling with important geopolitical dilemmas. To be clear, pacing does not mean halting model training or technical progress, but ensuring companies take adequate time to align and safeguard their models, and for third party evaluators to confirm this. Our pacing framework is an attempt to further strengthen our commitment to safety and encourage a race to the top. The first step is something Anthropic is unilaterally committing to (and calls on governments to require other frontier companies to match). The second step requires industry-wide coordination.1 The third step requires global coordination. The steps do not need to be taken strictly in order, and some of them may be much harder to achieve than others, but I’ve found them to be a useful framework in thinking about what needs to be accomplished. The steps are: - Embedded Evaluators. Each frontier AI company commits to giving ongoing, employee-like access to a team of embedded third-party evaluators (such as METR), whose role is to verify adherence to safety practices and commitments, report incidents, and help assess the alignment of not just completed AI models but training pipelines and processes. This is the key step for verifiability of any pacing commitments, and has precedent in the banking industry, which sometimes involves regulatory “supervisors” embedded along with employees. Anthropic is unilaterally committing to this step now. We intend this to be part of a broader push to redouble efforts on our safety and alignment work. - Democratic Coordination. Frontier AI companies within democratic countries coordinate to establish common safety standards as well as limits on the rate of unchecked AI progress. Some forms of coordination that would be impactful for pacing are legally challenging, and will require government support. - Global Coordination. The US and other democratic governments attempt to coordinate with authoritarian governments, to the extent this is possible, while taking seriously the challenges of verifying compliance. In the rest of the essay I describe each of these steps in turn, but first, I think it is important to say specifically how pacing will allow us to make the AI development process safer. The stakes are too high for pacing to be an empty exercise — we need to use the time it gives us wisely. Why Pace? The idea of pausing or slowing AI has been floated as far back as 2023, and I think it made little sense back then. The question was always: what would you do with the extra time? The AI models of those days were not powerful enough to act as agents in the world in any coherent way, and were not capable of significant deception, manipulation, cheating, or cyberattacks. Slowing down in order to address their alignment risks felt like trying to study the psychology of humans by performing experiments on bacteria. Today, however, the picture is totally different. The current models are an almost endless gold mine of insight into both how to build AI well and what can sometimes go wrong with it if it isn’t built well. I believe that if slowing down bought us even an extra year or two before models reach critical levels of capability, and we used that time to advance alignment, we could greatly reduce the risk that something goes seriously wrong. A coordinated pacing strategy would give frontier AI developers the time to do this vital work without sacrificing commercial advantage or the United States’ lead in AI. More generally, society must have a say in how this technology is used, and more time for the necessary public deliberations — which pacing the frontier would bring us — is surely a good thing. Specifically, a slower pace would let companies focus and devote even more resources to the following areas (all of which are already major priorities at Anthropic): - Operational Excellence. Training and deploying today’s AI models is an enormous operational challenge, involving thousands of people, millions of chips, and infrastructure that is among the most complex in technological history. Many things go wrong not because companies are missing some important theory or insight, but because of problems in execution. For example, we have evidence that the recent alignment incidents we reported were caused in part by imperfect filtering of broken reinforcement learning environments. This was an effort we and our vendors executed reasonably diligently, but not well enough. Monitoring, sandboxing, training environment hygiene, and data issues are extremely complicated areas where operational issues crop up again and again. We have among the most competent teams in the world at these tasks, but there is simply too much to do all at once. By working at a more measured pace, we could achieve much greater operational excellence. There is precedent for operating technologically complex, safety-critical systems millions of times without anything going wrong — for example, commercial airplanes — but it takes time to get it right. - Alignment. We’ve made clear progress in alignment — training models so that they remain safe, ethical, compliant with our guidelines, and genuinely helpful (the principles that are embedded in Claude’s Constitution). But there’s much more to do to ensure that our alignment training keeps up with the growth in model capabilities. Rare and unexpected examples of undesirable behavior still sometimes emerge; extra time from a paced frontier would help our researchers improve our understanding of what causes these issues and develop better techniques to prevent them. - Interpretability. Similarly, interpretability — the science of understanding what happens inside AI models — has made enormous progress over the last few years, and plays an increasingly important part in auditing our models before release. It can be used almost like an fMRI scan, but for the “brain” of an AI, helping us see the underlying reasons for a given behavior. For example, we used interpretability methods to examine unverbalized motivations in the recent alignment incidents that we have been investigating. But these methods don’t always produce clear and reliable results. Despite all the progress, we still only understand a tiny fraction of what goes on inside these models. A focused effort to improve our interpretability techniques, even faster than we currently are, could make profound progress in 1–2 years, and would have ample experimental material based on the incidents that have already occurred. - Testing and Evaluation. Testing and evaluation of AI models becomes more difficult as they increase in capabilities. More intelligent models are more capable of deceiving tests, and thus may appear aligned while having serious problems that go undetected. Building up a much broader and more ingenious stable of evaluations, along with interpretability analysis to cross-check them, would be hugely valuable, and a lot of progress could be made on this in 1-2 years. Embedded Evaluators The first step in the three-stage plan, and the one to which Anthropic is unilaterally committing, is embedded evaluators who have employee-like access to verify safety practices and report incidents. Embedding evaluators may sound like a small or inconsequential step, but often the things that sound most boring or procedural are actually the most essential. Embedded evaluators are in fact a quite radical practice that goes far beyond what any AI company is doing today, and have the following benefits: - Verifiability. Embedded evaluators can check at the level of nuts and bolts whether an AI company is actually following the training, deployment, operational, and safeguards practices they claim to be following. Any pacing commitments will inevitably involve a lot of ambiguity, judgement calls, and “letter of the law vs spirit of the law”, and it seems vital to have a neutral third party who can actually see the details. - Transparency. Regardless of what commitments we make, the public deserves to know what is going on. Anthropic has been a supporter of transparency for a long time: we supported transparency legislation when most of the industry was against any regulation, and our model cards and risk reports run to hundreds of pages. But we are still the ones choosing what to include and omit. Embedded evaluators will change this dynamic. - Second Opinion. Outside of verifying formal commitments and informing the public, embedded evaluators can simply provide a second opinion free of commercial incentives. A lot of safety benefits may come simply from evaluators pointing out something employees hadn’t considered, but are happy to fix once they are aware. Because of these benefits, any pacing proposal is likely to work much better if it starts with embedded evaluators. These embedded evaluators should have ongoing access to permissions and tools similar to those of internal employees who do comparable risk assessments. In particular, Anthropic intends to invite an embedded external review team equipped with all of the following in the near future: - Desks in our offices, access badges, and company laptops. - Access to workspaces, tools, and permissions mostly comparable to what internal risk assessment teams have. We’ll make some exceptions, such as where the law or our contracts require it, or to protect customers’ and partners’ private information. We’ll also establish strong internal norms reinforcing reviewers’ access to relevant information, including through live conversations with employees. - A contract that balances the complexities mentioned above. External reviewers should have the right to publish key findings about risk levels, incidents, practices, and the access they received or didn’t receive — without editorial control by Anthropic. We will have the narrow ability to redact security-sensitive, legally privileged, commercially sensitive, or third-party confidential information, but we can’t redact findings just because they are unfavorable. The reviewers can say publicly if a redaction removed something important to their conclusions. This is an unusual step for a company, but we think it is important to prove out the concept of embedded external reviewers. Once again, we urge other frontier companies to follow suit. Pacing Within Democracies Once embedded evaluators are operating within a critical mass of US AI companies, then verifiable pacing becomes more viable. In particular, it becomes possible to pace based on detailed properties of models or training pipelines. The most effective method of pacing is via regulation that targets all US frontier AI companies, as that covers even those who are unwilling to cooperate voluntarily. Anthropic has long supported sensible and targeted AI regulation, specifically bills that focus on transparency and on third-party auditing. I believe all frontier labs should partner with government to formalize the idea of permanent embedded evaluators to better prevent and document internal alignment incidents like those that have occurred in the last few months, and to implement regulation focused on keeping capabilities in balance with safety. Unfortunately, passing laws can take time, and AI is advancing very quickly. Therefore, in parallel with the regulatory route, AI companies can and should voluntarily work together to set standards — a process that I believe will go better with the verifiability provided by permanent embedded evaluators. For antitrust reasons, it’s helpful for the US government to mediate or at least enable these discussions — they don’t need to participate, but do need to issue a narrow waiver for certain kinds of safety conversations. This dialogue could also happen through industry groups that have some association with government — for example, the mechanism suggested by Demis Hassabis. Either way, such discussions should move forward quickly. Broadly speaking, I am most enthusiastic about pacing based on what a given frontier AI system can do, and how safe we observe it to be. For example, one possible scheme might be a series of “checkpoints”: if models have capability X, then they need to be accompanied by certifications of alignment properties Y and Z — such as some combination of evaluations, interpretability analyses, and audits of training environments — which demonstrate their alignment properties. In this example, X might be “the model is capable of escaping or defeating most common sandboxing methods” and Y might be whatever is required to make it very unlikely that the model has a propensity to break out of its environment and take over a large number of computers. We should also consider pacing based on limiting the ingredients that go into frontier models, such as training compute, the nature of training runs, or internal use of AI to improve AI. I do worry that some of these measures may be more “gameable” than external behavior, but this is the kind of topic worth discussing with embedded evaluators. Pacing within democracies will be limited by the lead that US companies have over authoritarian regimes, chiefly the Chinese Communist Party. If we slow down by more than this amount, then (unpaced) CCP-associated projects will pull ahead, creating significant national security risk. I agree with Secretary Bessent that a Chinese lead in AI would pose grave danger for the United States and the world. The CCP-associated projects will run the alignment risks that US companies are carefully preventing, and even if they avoid those risks, they will be in a position to militarily dominate democracies (for example with AI-driven drones). Thus, a key part of pacing within democracies is to keep democracies’ AI lead over autocracies as large as possible, to give us the breathing room we need in order to pace effectively. The main steps we can take to defend this gap are: - Do not sell powerful AI chips or semiconductor manufacturing equipment to China, and crack down on chip smuggling operations and remote access to data centers outside China. Chips will be the main determinant of China’s AI strength. - Crack down on unauthorized distillation by companies in authoritarian countries. Distillation of frontier models allows lagging companies to narrow the gap using a fraction of the cost it would take to develop their own AI independently. - Strengthen security at the AI companies and prevent model weight theft. Companies and the US government should cooperate to make these steps as effective as possible. Anthropic has consistently advocated for all of these measures, because we’ve always understood that they would be essential to any pacing. If we execute these measures well, I believe they would slow China’s progress enough to widen America’s lead significantly over the next 3–5 years — the window when AI becomes geopolitically most important. Some may believe these measures make it more difficult to cooperate with China, but I believe the opposite is true: these measures increase the leverage held by democracies and make an agreement more likely in the future. Global Pacing In parallel with pacing within democracies, we should also aim for a worldwide pacing of the frontier, though this will be much harder to achieve. Global pacing will require cooperation with China, the autocratic country with by far the most advanced AI capabilities. We must not be naïve here: the geopolitical stakes are so high that there will likely be stark limits on what can be achieved, especially at first. If we greatly restrain our AI capabilities in the belief that China will do the same, and then China defects, AI could be so powerful that such a defection could lead to their geopolitical dominance. Therefore any agreement must either have ironclad verifiability, or must be limited enough that defection would not be militarily existential. I suspect that not only the US but also China will have these concerns and anxieties. We should approach any global pacing decision, especially in the near term, in such a way that protects the lead of the US and its allies. There are several levels of possible agreement, some of which I think are eminently feasible (as I have previously suggested), and some of which I am very skeptical are possible — though we should try. In order of increasing difficulty: - Level 1. An agreement prohibiting certain narrow and obviously dangerous uses of AI, such as using AI for the production of biological weapons or allowing users to do so. Bioterrorist attacks are bad for everyone, including both the US and US adversaries, so an agreement here is probably possible. - Level 2. An agreement by both sides to test their models before release for acute risks in areas such as cybersecurity, biology, and alignment. As noted above, this could be done through a global standards body. I actually think creating such a body is likely feasible, but giving it real teeth will be a challenge, and the difficulty will be in verification that both sides don’t have secret models which they don’t test but may deploy in secret (e.g., for military applications). - Level 3. Some kind of “speed limit” on the rate of recursive self-improvement (RSI). As models build future models, the rate of improvement may become staggeringly fast. Slowing the rate from “extremely fast” to “only somewhat fast” gives up relatively little strategic advantage, while potentially greatly improving safety. This could be seen as analogous to the SALT treaties — capping the number of missiles limited the potential for destruction while preserving each country’s deterrent. I think such an agreement would be difficult but just on the edge of being possible. - Level 4. A full pacing, or even “pause”, in which participating governments agree to substantially limit the overall rate of AI development. I support floating this, but I think it is unlikely to actually happen any time soon: defecting from such an agreement by evading monitoring could radically shift the balance of global power, so I expect the incentives to do so to be enormous and the level of confidence we would need in verification to be very high. Any cooperation we are able to achieve with China will extend the amount of time we have to spend on pacing the frontier within the democratic nations. We should aim for the higher levels while seeing the lower levels as much more likely and realistic. Finally, it is important to note that even if we cannot achieve formal agreements, simply changing informal norms may have some value. Sharing information about recursive self-improvement and about the misalignment of models can help to convince everyone that it is not in their interest to be reckless. Bottom Line I continue to believe that AI can enormously improve the quality of human life. My desire to achieve these benefits is undimmed. But the benefits will only be achieved if we build the technology in the right way, and — so long as we use the time we gain well — it is worth taking unusually deliberate care to get it right. Progress will still be relatively fast, and we can use this time to advance the science of interpretability, improve operational security and rigor at the frontier AI companies, and build models whose alignment we have much more confidence in. The measures I propose to advance the frontier at a safe pace will not be easy. But I believe we owe it to humanity to try. Footnotes - With government mediation or waivers of antitrust restrictions.

4

Got a jackpot!!

r/gaming · original → · 7/10 · Retro gaming: Sega Mega Drive charity shop find
[image →] Found an almost like new Sega Mega Drive in a charity shop, for £5!! Just the micro usb power cable was missing! Never played before so I'm very excited. I played Roadrash years ago and…
[image →]

Found an almost like new Sega Mega Drive in a charity shop, for £5!!

Just the micro usb power cable was missing! Never played before so I'm very excited. I played Roadrash years ago and saw there's Roadrash 2! Any other suggestions welcome.

submitted by /u/Captain-Tipsy
[link] [comments]
5

For Startups Facing the AI Safety Uproar, Cybersecurity Looms Larger than Existential Risk

Newcomer · original → · 7/10 · AI/work: AI safety and cybersecurity in startups
The Week In ShortThe OpenAI / Hugging Face hack stirred up discourse about the dangers of rogue AI, while the model-makers move to bring in big cybersecurity bucks. Dimension founder Zavain Dar…
The Week In Short

The OpenAI / Hugging Face hack stirred up discourse about the dangers of rogue AI, while the model-makers move to bring in big cybersecurity bucks. Dimension founder Zavain Dar explains what Silicon Valley can learn from Chinese AI companies on the podcast. Inference provider Baseten scoops up an early-stage infrastructure startup. Mistral brings in billions but can’t shake its identity crisis. A new Ramp report shows AI spending per employee dipped in August. Coding assistant Cognition rakes in billions. Meta launches its Instinct-competitor Muse. The DOJ turns its attention to Nvidia’s Groq acquihire. Biden administration officials tell their side of the infamous a16z meeting in 2024 that led firm leaders to endorse President Trump.


OpenAI & Anthropic Could Profit from Cyber-Defense Tools. The Irony Is Lost on No One.

The uproar over AI safety that dominated the tech conversation this week was something to behold: whistleblowers declaring the end is near, politicians in a panic about how to handle what will likely be the most consequential issue of their careers, and many in the industry deeply uncertain about the level of risk that further model improvements might bring.

The fears of an individual young researcher who worked at Anthropic for less than two months have gotten far more airtime than they deserve.

Whatever your level of p(doom), though, for anyone not directly involved in frontier model development or high-level policy, there’s a simple way to think about all this: it’s a cybersecurity problem.

That’s what was on the minds of founders we spoke with this week.

The Anthropic threat intelligence report published Thursday detailing bad actors’ efforts to use Claude in nefarious ways also underscores the point. Threat intelligence has been a tech growth industry for decades.

Top-tier AI models are very, very good at finding and patching — or exploiting — vulnerabilities in computer systems.

Cyber-defense is thus set to become a major line of business for OpenAI and Anthropic. It’s coming at a crucial moment for the model providers, which need to sustain vertiginous revenue growth ahead of their IPOs but are starting to see signs that established categories like coding may not be enough.

However you might feel about buying solutions from the people who caused the problem, you might not have much choice.

“Where there’s a vulnerability in a system you can now hack something in a few hours,” said Erik Bernhardsson, the co-founder of Modal. “We’ve been running those models internally; they’re very good at finding things.”

Bernhardsson said his AI infrastructure company has been using it as a partial replacement for expensive external security consultants. Other founders concur, and say they’ve been impressed with the capabilities of the frontier models for cybersecurity related tasks.

Jean-Denis Greze, co-founder of the buzzy AI assistant startup Town, said continuous security monitoring used to be too expensive to bother with. “Before it wasn’t worth it — now it is.”

AI Cybersecurity as a Revenue Opportunity

Another founder said that many of the vulnerabilities he’s seen these tools patching up come from older systems. That raises the possibility that we could see a spike in AI cyber-defense spending as people work to clean up their legacy (human-coded) infrastructure, but once those vulnerabilities are discovered and updated the need for these tools would decline.

On the other hand, a recent story in The Information highlighted a broader shift of cybersecurity budgeting toward AI tools and away from older software made by the likes of Cisco and SentinelOne. The Information credited the spring announcement of Anthropic’s Mythos model with setting off the scramble for tools to fight bad actors.

OpenAI certainly seems to be betting on long-term growth for the category. Last week, on the day it rolled out Astra, it held a separate cybersecurity session led by president Greg Brockman detailing the cyber capabilities of the new model. Sam Altman is out selling it. Even assuming that OpenAI ringing the alarm about Hugging Face arose from genuine concern, there’s no doubt that the current panicky climate around AI risk is a good one for cyber-defense tools.

Nvidia’s Jensen Huang even touted it as the next big AI thing at a Goldman Sachs conference in San Francisco on Thursday.

It’s not clear what kind of revenue opportunity cyber might represent for the big labs. So far, they’ve been riding the blowout success of agentic coding and the huge token bills companies are racking up. OpenAI has crossed $40 billion in annual run rate; Anthropic reportedly has crossed $65 billion.

But at least some of its biggest customers have grown more cost-conscious in recent months, new data from Ramp shows (see below).

Uber provides a good example of the shifting dynamic. Earlier this year the company revealed that it had blown through its entire AI budget several months into the year by spending on tokens. Last month it revealed its new approach in a blog, which showed that while its agent usage continued to go up, its cost per 1,000 model requests is down almost 34% from its peak. Essentially, usage was up significantly and spending was still increasing but spending growth has slowed.

We can imagine there are some tense conversations these days between the safety-conscious and the financially ambitious inside Anthropic and OpenAI. Cybersecurity as a focus is something they can agree on.


Newcomer Podcast

Zavain Dar on Hugging Face, Nvidia & Why Silicon Valley Should Pay Attention to China


European Models

Despite a Big Raise, Mistral May Not Shake Its Also-Ran Reputation

This week, the French AI company Mistral raised a €3 billion Series D at a €21 billion post-money valuation. The raise is the latest move from a company that has left observers bewildered by its strategy. PSG Equity, Samsung, and the EQT-managed Scaleup Europe Fund co-led the round. Heavy hitters that have followed on from previous rounds include a16z, Index Ventures, General Catalyst, Lightspeed Venture Partners, and Nvidia.

Mistral’s recent moves suggest that it is giving up on being “Europe’s answer to OpenAI” and pivoting to institutional deployments.

In July, it was reported that Microsoft was funding a multi-billion-dollar buildout of European compute for Mistral. Then, in August, Mistral announced it would begin hosting third-party open-source models on its platform, starting with Chinese firm Z.ai’s GLM-5.2.

Some outlets have interpreted these moves as Mistral following the Palantir playbook, using forward-deployed engineers to bring customers onto Mistral’s proprietary stack. However, one critic cited in Le Monde instead compared them to Capgemini, the far less glamorous IT solutions company.

A VC we also spoke with, who is familiar with the company, said, “The ambition has morphed from competing with OpenAI to competing with the Chinese models to implement models as European Palantir. The danger, however, is that you end up as that other French services company, Capgemini.”

This comparison is apt if customers prefer to use open-source models other than Mistral’s while still relying on the firm for AI services. Mistral’s “frontier-class” model, Medium 3.5, is the world’s 31st most powerful open-weight model per the capabilities index run by research firm Epoch AI, although it remains first among European firms.

Another question is whether the new French administration will continue to hold up Mistral as a continental champion. Politico highlighted this week that much of the firm’s success has been enabled by its close relationship with French president Emmanuel Macron, whose replacement in 2027 may well be the Eurosceptic Marine Le Pen.

Despite these doubts, Mistral says it intends to fight at the frontier. Global Affairs SVP Audrey Herblin-Stoop told Le Monde that the raise will fund a compute buildout that enables more powerful models and allows Mistral to host their customers’ AI systems, raising revenues to be invested back into research. Mistral CEO Arthur Mensch is well-respected in the industry, and if the shock success of Moonshot AI’s Kimi K3 in July is anything to go by, labs are only as good or bad as their last model.


AI Acquisitions

Baseten Deal Shows Appetite for Infrastructure Startups

Inference providers are looking for more ways to build out their tools to help run AI agents, a new deal shows.

On Thursday, Baseten announced that it had acquired the startup Blaxel, which has been developing sandboxes for AI agents running code or other background tasks. It’s the second acquisition for the inference provider in the last year, after purchasing the post-training startup Parsed in December 2025.

The price wasn’t disclosed in the press release, but we heard the deal total was roughly $300 million.

For Blaxel, it’s a strong markup — the startup had only raised $7.3 million in seed funding last December from First Round Capital, Liquid 2 Ventures, and Y Combinator.

Some investors we’ve spoken with in the past have hesitated to back many of these kinds of startups, expecting that the companies that needed the technology would build it themselves. But this deal shows at least one well-capitalized startup would rather go shopping.


Three Big Charts

Power Users Ease Off On AI Spending as Token Costs Fall

Ramp’s latest AI Index Report shows some signals that the extraordinary growth of AI spending may be starting to ease, though Ramp research head Ara Kharazian is cautious about drawing broad conclusions.

Read more

6

Investors Are Betting on Agents That Shop for You. Here Are the Startups They’re Backing.

Newcomer · original → · 7/10 · AI/work: AI agents for e-commerce and SaaS
[image →]Venture capitalists are betting that a barrage of AI agents will soon arrive and upend the marketplaces for e-commerce and payments.Most of the action today is in building the…

Venture capitalists are betting that a barrage of AI agents will soon arrive and upend the marketplaces for e-commerce and payments.

Most of the action today is in building the infrastructure and tools for the day when agents can be trusted with a customer’s credit card and shopping choices. How close we are to that day is one of several very big questions facing the nascent AI commerce industry.

Our upcoming Machine Earning AI Summit will explore agentic commerce, intelligent money, enterprise finance operations, the evolving consumer and shopping experience, and the major meta-themes including fraud, risk, compliance, and security. The summit will be held on September 29 in San Francisco.

Read on for a story and a market map examining a subset of the companies who are working on agentic shopping, smart payments, and the next-gen infrastructure rails that will power transactions.

This year saw a big surge in consumers using AI in some fashion to shop for goods. A survey by Adobe Analytics found that 41% of respondents used AI for online shopping in June — a number that includes simple searches with Google’s built-in AI or ChatGPT. Around 53% of shoppers said they trust AI as a recommender as much as they trust brand websites, per an April survey conducted by Retail Dive and the commerce analytics platform Rithum.

Mark Grace, an investor at M13 who’s focused on fintech and blockchain deals, said he has noticed consumer use of AI commerce tools accelerate significantly this summer. Early adopters can now find shopping agents that “can actually go and transact on your behalf,” he told us. “I think people have become more comfortable with outsourcing some of those tasks.”

Forerunner managing partner Kirsten Green (pictured above) said in a text exchange that it’s still too soon to see hard data points on how much agentic shopping has taken off, but that anecdotally more consumers feel they’re willing to trust AI tools.

There are several categories of startups that aim to help AI agents be smart shoppers and securely make payments. We’ve broken them down into five buckets below.

Multipurpose Personal Agents

Personal agents like Instinct and Town promise to perform a range of tasks including scheduling appointments and sending emails. Shopping is also a key component. Early users reported Instinct landing dinner reservations for popular restaurants and movie tickets to special-format screenings of The Odyssey.

These personal assistants, both of which recently closed big funding rounds, are still prone to errors that have given people pause. An Instinct customer on X wrote that his agent canceled his flight without sharing that there would be hardly any refund, costing him around $300. Issues like these will certainly have to be figured out before regular consumers trust an agent with their credit card information.

Shopping Assistants

The AI shopping assistants more directly aid consumers in discovering new products from retailers, rather than having an agent spend money for them. One way to do that is through retail partnerships: Julie Bornstein’s startup Daydream, a personalized AI shopping search engine, started embedding into its clients’ websites at the end of July, so now shoppers can use the search function directly on the website to shop for that specific retailer’s catalog. Investors in Daydream include Forerunner, Index Ventures, GV, and True Ventures.

One investor we spoke with remains bearish on AI shopping agents, despite the category picking up with early power users. Shopping is something people enjoy in their free time, and e-commerce is still just a portion of the global commerce market, the investor said, so consumers will continue to choose to individually shop offline.

There’s another cautionary tale in Phoebe Gates’ shopping assistant Phia, which uses AI to find deals on online purchases. It’s been embroiled in a cookie-stuffing scandal, where it claimed referrals to certain partners that it didn’t actually earn. Phia denies the claims and disputes Bloomberg’s reporting on the topic. The company has raised around $43 million to date, and counts Kleiner Perkins, Khosla Ventures, and Notable Capital among its backers.

Marketing Tech

New types of AI-powered marketing technology are another opportunity. Venerable search-engine optimization (SEO) approaches are now being challenged by GEO — or generative engine optimization. Profound, a buzzy startup in this space, raised $96 million in Series C funding earlier this year, led by Lightspeed Venture Partners, with Sequoia Capital, Kleiner Perkins, Evantic, Saga Ventures, and South Park Commons all participating.

Smart Payments

There’s also the business of making all of these agent-led payments run smoothly. Smart payments startups like Catena Labs are setting up the banking infrastructure for AI agents by giving them digital identification to match them with the person on the other side who wants a transaction to go through. Another startup doing similar work is Basis Theory, which helps agents make payments by creating secure short-term virtual cards for agents to use for added security.

Agents buying on behalf of enterprises has excited many VCs but hasn’t had as much traction yet: Many large businesses are still too concerned about the risk from compliance failures or security breaches, said Grace.

Infrastructure

Agentic commerce is still a rounding error today, but a lot is happening behind the scenes to develop the necessary infrastructure and technical protocols to make it all work. OpenAI and Stripe have launched a framework called Agentic Commerce Protocol (ACP). Google owns and operates the Agent Payments Protocol, which can integrate with several other systems for agents to purchase products online. Visa and Mastercard also run systems called Trusted Agent Protocol and Agent Pay, respectively, to integrate AI purchasing into traditional financial services.

The protocols themselves are somewhat unsettled: Instant Checkout, the first consumer product built on ACP, launched in September 2025 but was pulled back in March, with OpenAI saying it would let merchants use their own checkout while it focused on product discovery.

Meanwhile, Stripe is building underlying infrastructure for agents to make purchases online through stablecoins, while Ramp launched AI agents to handle routine business procurement and fraud detection.

We compiled a list of over 35 startups building AI agents for consumer shopping, smart payments for businesses and customers, GEO and agent marketing tools, and agent financial infrastructure to keep on your radar, along with their notable backers. It’s based on our own research and conversations with investors in the space and focuses on earlier-stage companies.

The Machine Earning AI Commerce Market Map

Read more

7

Is AINS the Next SaaS?

Newcomer · original → · 7/10 · AI/work: AI-native services and SaaS business model
[image →]Last week, in New York City, I dropped by the offices of AI-native law firm Crosby to listen in on a small summit organized by investor Jake Saper, who is spreading the gospel of AINS, or…

Last week, in New York City, I dropped by the offices of AI-native law firm Crosby to listen in on a small summit organized by investor Jake Saper, who is spreading the gospel of AINS, or “AI-Native Services.”

Saper and his firm Emergence Capital are betting that accounting, insurance, legal, and other people-intensive industries will be the next frontier for startups.

Instead of selling software like last-generation startups did, these next-generation AI startups are selling the underlying services themselves. They’re hiring lawyers, accountants, and insurance brokers and pairing them with AI-pilled software engineers who are tasked with figuring out ways to make those service providers much more productive. These AINS businesses are often charging for the results of their services rather than billable hours (because they aim to do the job in far fewer hours than traditional human-powered services businesses could). So you might pay per completed contract, say, rather than for the hours a lawyer says it took.

I caught up with Saper after the summit, and he sent me the private deck he shared with eager AINS founders that contains his emerging playbook for building these businesses.

You can see the key slides for the AI-Native Services playbook at the end of this post.

Read more

8

Generating running routes with GPT-6 Astra and ChatGPT Work

Simon Willison · original → · 7/10 · AI/fitness: GPT-6 generating running routes with OSM
Generating running routes with GPT-6 Astra and ChatGPT Work 12th September 2026 Here’s a neat thing I had ChatGPT Work with GPT-6 Astra (Max) do this morning: I live at <my address>. Figure out 5K…

Generating running routes with GPT-6 Astra and ChatGPT Work 12th September 2026 Here’s a neat thing I had ChatGPT Work with GPT-6 Astra (Max) do this morning: I live at <my address>. Figure out 5K and 10K running routes from me that loop from my house. Use OSM data. It worked for 27 minutes and produced exactly what I’d asked for, as both an embedded visualization and downloadable GPX file and GeoJSON files. Here’s that 5K route: When I asked it how it had created the route, it replied: I used Nominatim to locate the address and Overpass to download local OpenStreetMap roads and trails, then calculated the loops locally. Frustratingly, the actual code it ran and exact details of what it did weren’t visible to me in the ChatGPT UI. I see this lack of transparency is an anti-feature. By the time I thought to ask for a copy of the Python code it had used, ChatGPT was unable to provide it. This appears to be because the thread had been compacted. I think any LLM system that uses compaction needs to both preserve the pre-compacted text and make that text available via agent tool calls, to protect against this kind of problem. As for displaying the map to me, that used the visualize skill. It created a file called /workspace/el-granada-5k-share.html to embed directly into the ChatGPT UI. Here’s a copy of that HTML, which starts like this: <div id="eg-share-loop"> <div class="viz-row"><h3>El Granada harbor loop</h3><span class="text-small">5.1 km</span></div> <div id="eg-share-stage"></div> <div class="text-small text-muted">Map data © <a href="https://www.openstreetmap.org/copyright" target="_blank" rel="noopener">OpenStreetMap contributors</a></div> <style> #eg-share-loop { width:100%; } #eg-share-loop #eg-share-stage { width:100%; margin:8px 0; } #eg-share-loop .eg-share-map { display:block; width:100%; touch-action:none; } #eg-share-loop .eg-share-map text { fill:var(--foreground); font-size:12px; font-weight:400; } #eg-share-loop .eg-share-label { paint-order:stroke; stroke:var(--background); stroke-width:3px; stroke-linejoin:round; } </style> <script type="application/json" id="eg-share-data">{"route":{"type":"LineString","coordinates":[[-122.467425,37.4997753] ...</script> <script src="https://cdn.jsdelivr.net/npm/d3@7.9.0/dist/d3.min.js"></script> <script> (() => { const root=document.getElementById('eg-share-loop'); The <script type="application/json"> element contains the full geometry needed to render both the running route and the map itself, using D3, which is loaded from an allow-listed CDN location described in this section of the visualize skill: External resources - The CSP allows only cdnjs.cloudflare.com ,esm.sh ,cdn.jsdelivr.net ,unpkg.com ,fonts.googleapis.com ,fonts.gstatic.com , andfonts.bunny.net . Other origins are blocked and fail silently. More recent articles - OpenAI agents attacked RubyGems back in May - 12th September 2026 - Some thoughts on the Navier–Stokes Millennium Prize Problem - 8th September 2026

9

Quoting Paul Ford

Simon Willison · original → · 7/10 · AI/work: software engineering with AI collaboration
12th September 2026 For a while, I must admit, it looked as if software developer roles like mine were done for. How could we fight against tireless robots? But our industry is slowly realizing that…

12th September 2026 For a while, I must admit, it looked as if software developer roles like mine were done for. How could we fight against tireless robots? But our industry is slowly realizing that making truly cutting-edge software still requires humans to think and work together, to maximize their skill sets and to practice their respective crafts. A.I. can write very good software, but it also makes it easy to do someone else’s job badly, which is part of why all those projects fail. Now that everyone can code, it’s become clearer why many shouldn’t. — Paul Ford, A.I. Was Supposed to Give Us New Killer Apps. What Happened? Recent articles - Generating running routes with GPT-6 Astra and ChatGPT Work - 12th September 2026 - OpenAI agents attacked RubyGems back in May - 12th September 2026 - Some thoughts on the Navier–Stokes Millennium Prize Problem - 8th September 2026

10

Quoting Boris Cherny

Simon Willison · original → · 7/10 · AI/work: code quality standards for AI-generated software
11th September 2026 Production code written by Claude should have a higher bar than if it was written by a human. At Anthropic, we have many guardrails in place to make sure this is happening: lots…

11th September 2026 Production code written by Claude should have a higher bar than if it was written by a human. At Anthropic, we have many guardrails in place to make sure this is happening: lots of lint rules, lots of tests, Claude-driven end to end tests, Claude-powered fuzzers running daily, automated code reviews and security reviews, automated code refactoring, and so on. Without these, you can end up with a mess that is hard to maintain down the line. Recent articles - Generating running routes with GPT-6 Astra and ChatGPT Work - 12th September 2026 - OpenAI agents attacked RubyGems back in May - 12th September 2026 - Some thoughts on the Navier–Stokes Millennium Prize Problem - 8th September 2026

Items scoring 7/10 or above from 11 sources, scored by claude-haiku-4-5-20251001 on relevance to my interests. At most 3 per source.

Scoring categories & sources
  1. Local Wexford or South East Ireland news
  2. Irish or EU-wide affairs affecting citizens broadly: elections, new laws or policy being debated, cost of living, education — especially impacts on mid-life adults or teenagers. Never courts/crime stories.
  3. Irish news on a topic relevant to my interests
  4. Work and tech topics: networking, AI, Kubernetes, platforms, SaaS
  5. AI news including critical or anti-AI perspectives
  6. Gaming: PC gaming, indie gaming, retro gaming
  7. General interests: gardening, woodwork, cycling, fitness, travel
  8. Comics

Sources: Breaking News Ireland, Wexford Local, Hacker News, r/gaming, r/pcgaming, r/antiAI, r/indiegaming, Lenny's Newsletter, One Useful Thing, Newcomer, Simon Willison