daily

2026-08-08
1

Managing AI Coding Costs at Scale

Hacker News · original → · 8/10 · AI/work: managing AI coding costs at scale, platforms/SaaS relevant
by Patrick Wendell, Akshat Bhatia, Vinay Gaba, Erich Elsen and Ivan Zhou AI coding tools deliver immense value: at Databricks, agentic coding has measurably improved every velocity metric we track…

by Patrick Wendell, Akshat Bhatia, Vinay Gaba, Erich Elsen and Ivan Zhou AI coding tools deliver immense value: at Databricks, agentic coding has measurably improved every velocity metric we track and, in some teams, driven an order-of-magnitude gains in output. But nearly every company deploying AI tools at scale has hit the same wall: exponentially growing costs. That curve is unsustainable - left unchecked it will eventually overtake revenue. The spend explosion has left enterprises in a paradoxical situation: on the one hand, desiring to maximally push AI transformation and put powerful tools in the hands of employees, and on the other hand, having to reconcile with an aggregate cost profile that threatens to undermine or even reverse the very efficiency gains AI provides. Fortunately, several of the earliest large-scale adopters have converged on a set of approaches that solve this puzzle, achieving a “dual mandate”: (a) providing broad access to AI tooling, with minimal friction, and (b) keeping aggregate costs inside of a roughly fixed envelope per user. This post outlines proven cost management techniques, based on our experience at Databricks and conversations with several other digital-native companies, including Stripe, Coinbase, Uber, and Ramp. The table below summarizes current techniques and associated savings; the numbers are directional, based on an informal survey of development teams: Some of these techniques can be easily implemented with software many companies already use. Others require new infrastructure, particularly techniques that modify end-user clients or shift traffic across models. At Databricks, we’ve open sourced or made freely available our key infrastructure components: an end user meta-harness (Omnigent) and our AI Gateway (Unity AI Gateway). For completeness, this post also covers software used by other companies we spoke with. The single greatest cost lever in moving coding spend to more efficient models as they are released. This point bears some discussion, as the simple explanation of "cheaper models” in fact hides a nuanced relationship between model cost and quality. Colloquially, the term frontier model means “the highest intelligence model,” and frontier labs largely focus on advancing peak intelligence. Frontier models can now solve novel problems in math or cybersecurity. But when AI is deployed at scale, a different type of frontier matters more: the efficiency frontier. The efficiency frontier is defined by the set of models that have the best price point for a given level of intelligence. Most day-to-day coding doesn't require mathematical proofs or novel security insights, so what matters in aggregate is the cost of models that meet the quality bar for typical software engineering work. This "efficiency frontier” is advancing far faster than the intelligence frontier, with new models being released almost weekly that present better intelligence-per-unit-price than prior models. Rapidly adopting newer, more efficient models delivers the largest cost wins of any technique. But to capture those gains, a company first needs to know which models actually beat its incumbents. This can be difficult because public benchmarks do a poor job of indicating real-world performance on coding tasks. To size up new models, many companies have built automated evaluations that they believe are more representative of their internal development mix. Databricks recently published an example of such a benchmark, in which we observed highly competitive price/performance for GLM models. That benchmark led us to roll GLM out to developers internally. Often, new models do not advance the efficiency frontier,and evaluations frequently produce negative results: Stripe found that Opus 4.7 did not meaningfully improve quality over Opus 4.6, while increasing cost. They therefore declined to make Opus 4.7 available internally. Databricks saw similar cost regressions when comparing Opus 5.0 to 4.8. Since the biggest wins come from switching to new models, adopting end user tooling that allows for model flexibility is becoming a critical component of keeping costs down. The tool most commonly used in concern with a particular model is called harness. Proprietary frontier models are increasingly co-designed to work well with specific harnesses, meaning certain harnesses “work better” with certain models. If a company wants to preserve model independence there are roughly two approaches: Ask users to switch harnesses. One approach is to provide developers with a set of harnesses (Claude Code, Codex, or Cursor) and then ask them to switch between harnesses when a company wants to migrate spend to lower cost models. This lets users work in their preferred harness when possible, but the downside of this approach is that switching costs for an individual developer can be high. If switching costs become too high, the harness itself becomes a de facto lock-in to a model family, limiting the ability to move spend to more competitive models. Use a meta-harness. A new and increasingly popular approach is to use a meta-harness that surfaces a common user experience to developers while dispatching requests to underlying harnesses (both proprietary and open source). This approach allows both model/harness independence while also reducing developer switching costs. At Databricks, this is the default mode for developers who leverage Omnigent. Some companies we talked to have built custom internal meta-harnesses that integrate with their development toolchain. Instead of asking users to choose task-appropriate models themselves, a growing body of research suggests that automatic model and tool selection may further squeeze efficiency out of agentic coding workflows. Routing approaches roughly fall into three categories: It may be surprising that this entire article did not start and end with “Give users a monthly budget and be done with it.” Hard budgets, where usage is entirely cut off at a specific spend threshold, are often used only as a last resort option in every company we spoke with. There are two reasons that hard token budgets are not particularly effective for AI spend management: First, if a developer hits their budget ceiling, cutting off further access to AI tools would be debilitating to productivity. Neither the company or employee actually wants that outcome. Second, at least some of the “high spending” users are in fact those who have achieved monumental efficiency gains with AI and are producing immense output. Discouraging those users is self-defeating. Instead of a hard user spending cap, most companies are adopting a more nuanced and progressive approach that focuses on visibility for end users and increased degrees of friction as spend increases. Visibility: Every company we spoke with had a mechanism to provide near-instantaneous feedback to users on their ongoing spend, with many also offering specific tips or insights on how to reduce spend by using less expensive models. It is important that users be able to see their spend across all tools, since they may want to influence their choice of tool where they get the highest ROI. A developer dashboard at Databricks showing active spend When a user types a relatively simple request into an AI coding agent (such as “Please investigate and fix this bug.”), that agent subsequently gathers massive amounts of relevant context, invokes a large number of tools, searches through the codebase, and integrates skills or system information provided by the company. By the time costly LLM inference occurs, the user's initial statement accounts for only a negligible fraction of the data fed into the AI system, meaning costs are dominated by context the user did not explicitly include. Techniques in reducing context bloat are still new, but several promising approaches are being explored, such as: When contexts get large, prompt caching also plays a meaningful role in overall performance. Both proprietary and open source LLMs have settings that allow you to enable prompt caching and tune how long the cache is stored. Cache writes cost money, but cached reads can drastically reduce per-inference cost. This trade-off is dependent on a company’s specific workload, so hand-tuning of default cache settings to increase overall cache hit rate can have drastic improvements to overall cost. At Databricks, relatively simple tuning of our harness and caching settings led to an almost 50% reduction in the number of generated tokens and associated costs, with no observed quality degradation for developers. We continue to explore techniques in this area and think meaningful additional optimization remains possible. A drastic reduction in tokens per session by eliminating extraneous inference calls and reducing cache writes. The techniques above had many implicit technical requirements: To rapidly take advantage of new models, companies must have a central location where the “model menu” is managed, and end-users must have a toolchain that supports model mixing. To provide budget visibility across multiple AI tools, a unified cost observability capability must exist. To manage context bloat, companies need a way to observe typical toolcall outputs and enforce compression or compaction. These needs are collectively being solved by a new class of infrastructure software, best described as an AI Gateway. An AI gateway is a central location where all of the following occur: At Databricks, we rely heavily on Unity AI Gateway for all of these capabilities. The exponential growth of AI coding costs is not an inevitability, it's a solvable engineering and governance problem. Companies that have tamed it share a common playbook: relentlessly chase the efficiency frontier rather than the intelligence frontier, adopt tooling that preserves model flexibility, route work intelligently to the cheapest capable model, replace hard budgets with visibility and progressive friction, and cut the token overhead that dominates real-world spend. None of these techniques requires sacrificing the productivity gains that made AI adoption worthwhile in the first place; together, they let organizations satisfy the dual mandate of broad, low-friction access within a predictable cost envelope. A set of new infrastructure abstractions is emerging to give companies the tools to manage their costs. At Databricks, we’ve released the key components in our cost management stack as open source or free software products: Our Unity AI Gateway for central management and Omnigent for developer tooling. Thousands of companies use these components every day. We invite more companies to share findings and compare techniques as this technology landscape rapidly evolves. Acknowledgements: Thank you to infrastructure leaders at Uber, Stripe, Coinbase, and Ramp who provided commentary and reviews of this article. Thank you to Thrive Capital for feedback on an early draft of this article. Subscribe to our blog and get the latest posts delivered to your inbox.

2

Oracle bans AI-generated code from OpenJDK

Hacker News · original → · 8/10 · AI: Oracle bans AI-generated code from OpenJDK, critical AI perspective
Oracle bans AI-generated code from OpenJDK despite Ellison's claim 'Oracle isn't writing' its own code ● 5 days agoOracle has banned AI-generated code from OpenJDK contributions, citing safety,…

Oracle bans AI-generated code from OpenJDK despite Ellison's claim 'Oracle isn't writing' its own code ● 5 days agoOracle has banned AI-generated code from OpenJDK contributions, citing safety, security, and intellectual property risks. The open-source Java project steward said developers can use LLMs privately for debugging and reviewing code but cannot submit AI-generated material to repositories, pull requests, or other project channels. The policy contrasts sharply with Oracle's internal practices. Co-founder Larry Ellison recently declared that AI models now write Oracle's code, whilst co-CEO Mike Sicilia credited AI tools with enabling smaller engineering teams to deliver faster. Oracle is investing $70 billion this year in datacentre expansion. The spending spree prompted credit agency S&P to downgrade Oracle's rating to BBB-, one notch above junk status, citing uncertain returns on investment. Source: theregister.comOracle Austin, United States70

3

Fleadh visitor numbers to Belfast expected pass one million by Friday

Breaking News Ireland · original → · 7/10 · Irish cultural event: Fleadh Cheoil success in Belfast, tourism impact
The success of the first Fleadh Cheoil na hÉireann in Belfast has been hailed as visitor numbers are expected to pass one million by Friday night. It surpasses the initial predictions of 800,000…

The success of the first Fleadh Cheoil na hÉireann in Belfast has been hailed as visitor numbers are expected to pass one million by Friday night. It surpasses the initial predictions of 800,000 visitors for the whole week of the world’s largest celebration of Irish culture and music. Stormont Economy Minister Caoimhe Archibald said her officials are also expecting the economic boost to be greater than the previously predicted £53 million figure. President Catherine Connolly, Taoiseach Micheál Martin, Secretary of State Chris Bryant, First Minister Michelle O’Neill and deputy First Minister Emma Little-Pengelly are among the hundreds of thousands in the Northern Ireland capital. It has included performances from Loyalist marching bands and minority ethnic groups as well as traditional Irish music and dance. Friday saw the start of the Fleadh competitions with more than 5,000 competitors from 22 countries vying for All-Ireland titles in music, dancing and sean nós singing. Initial estimates of visitor numbers to the event, which was expected to attract 800,000 across the weekend indicate that was already surpassed by Thursday. Some 176,462 visitors were recorded on Sunday for the opening ceremony – an increase of 95% on the same day last year. There were 165,077 visitors on Monday, 162,015 on Tuesday and 169,296 on Wednesday before jumping to 227,170 on Thursday. The past five days have brought huge energy, talent and craic into Belfast, showing us the potential of the city to attract events with a worldwide scope. Martina Connolly, CEO of Belfast One Business Improvement District (BID), said they expect the overall visitor number to pass one million by the end of Friday. She said the numbers have been recorded by footfall cameras at six locations across one square mile of the city including Ann Street, Castle Street, Donegall Place, Royal Avenue, Corn Market and Fountain Street “Since the beginning of Fleadh Cheoil na hÉireann, our Belfast One footfall cameras have recorded 900,020 visitors in the city centre up to Thursday 6th August, exceeding expectations and putting numbers on track to pass the one million mark in our area by the end of today, the first of the three biggest days for the Fleadh as the main competitions kick off and even more visitors arrive to soak up the atmosphere,” she said. “The busiest day so far, Thursday, saw 227,170 visitors recorded within our remit alone, which covers many of the key areas where the festival is taking place.” She added: “The past five days have brought huge energy, talent and craic into Belfast, showing us the potential of the city to attract events with a worldwide scope, which has been wonderful to see for organisers, performers, businesses and visitors alike.”

4

Senseless vandalism at ancient burial ground

Wexford Local · original → · 7/10 · Local Wexford: Monachán Graveyard vandalism, community concern
[image →]Aontú Cllr Jim Codd is “simply horrified” at criminal damage to the ancient Monachán Graveyard at Taghmon and Gardaí are appealing for information that would assist their investigation into…
[image →]
Aontú Cllr Jim Codd is “simply horrified” at criminal damage to the ancient Monachán Graveyard at Taghmon and Gardaí are appealing for information that would assist their investigation into the matter. (Pic; Cllr Jim Codd).

By Dan Walsh

Wexford Gardaí are seeking assistance with their investigation into damage to several headstones and a stone wall at Monachán Graveyard in Taghmon between 20 April and 27 July 2026.

The vandalism of at least 15 graves in the centuries-old cemetery has attracted national media coverage.

At some point in recent weeks, vandals are alleged to have broken and smashed the surrounds of graves, knocked over and broken headstones and pulled down a cast iron cross from a wall at Monachán Graveyard.

Locals were enraged by the discovery, blasting it as “disrespectful” and “absolutely disgusting” and branding those responsible as “the lowest of the low”.

In a social media post, local Aontú Cllr Jim Codd said; “This is not the actions of children. Two men would be busy with a sledgehammer to complete this attack on our Christian burial grounds.”

He said he was “simply horrified” by the attack.

Gardaí are appealing to anyone who may have witnessed any suspicious or unusual activity in the area during this time period to come forward.

Gardaí are also seeking road users, residents and businesses who may have relevant CCTV, dashcam or other video footage from the area to make that footage available to the investigating Gardaí.

Anyone with information is asked to contact Wexford Garda Station on 053 916 5200, the Garda Confidential Line on 1800 666 111, or any Garda Station.

Investigations are ongoing.

5

DeepSeek V4 Flash 0731

Hacker News · original → · 7/10 · AI: DeepSeek V4 Flash model benchmarks and performance
DeepSeek V4 Flash 0731 DeepSeek·Jul 31, 2026·3 reasoning variants At max effort, DeepSeek V4 Flash 0731 scores 89.0% on ARC-AGI-1 Semi-Private at $0.02 per task and 61.4% on ARC-AGI-2 Semi-Private…

DeepSeek V4 Flash 0731 DeepSeek·Jul 31, 2026·3 reasoning variants At max effort, DeepSeek V4 Flash 0731 scores 89.0% on ARC-AGI-1 Semi-Private at $0.02 per task and 61.4% on ARC-AGI-2 Semi-Private at $0.04 per task. ARC-AGI 2 leaderboard DeepSeek V4 Flash 0731 Verified scores | Variant | ARC-AGI-1 | ARC-AGI-2 | ARC-AGI-3 | |---|---|---|---| | Max | 89.0% | 61.4% | — | | High | 87.0% | 56.0% | — | | Low | 84.0% | 46.0% | — | Tasks & environments Pass/fail per reasoning level across each benchmark. ARC-AGI-2 Public Eval 120 tasks| Task | Max | High | Low | |---|---|---|---| | ✗ | ✓ | ✗ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✗ | ✗ | ✗ | | | ✗ | ✓ | ✗ | | | ✗ | ✗ | ✗ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✗ | ✗ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✗ | | | ✗ | ✗ | ✗ | | | ✗ | ✗ | ✗ | | | ✓ | ✓ | ✗ | | | ✓ | ✓ | ✓ | | | ✗ | ✗ | ✗ | | | ✗ | ✓ | ✗ | | | ✓ | ✗ | ✗ | | | ✓ | ✓ | ✗ | | | ✗ | ✗ | ✗ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✗ | | | ✗ | ✗ | ✗ | | | ✓ | ✓ | ✓ | | | ✓ | ✗ | ✗ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✗ | ✗ | ✗ | | | ✗ | ✗ | ✗ | | | ✓ | ✗ | ✗ | | | ✓ | ✓ | ✗ | | | ✗ | ✗ | ✗ | | | ✓ | ✓ | ✓ | | | ✓ | ✗ | ✓ | | | ✗ | ✓ | ✓ | | | ✗ | ✗ | ✗ | | | ✗ | ✗ | ✗ | | | ✗ | ✗ | ✗ | | | ✓ | ✓ | ✓ | | | ✗ | ✗ | ✗ | | | ✗ | ✗ | ✗ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✗ | ✗ | | | ✗ | ✗ | ✗ | | | ✗ | ✓ | ✗ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✗ | ✗ | | | ✗ | ✗ | ✗ | | | ✓ | ✓ | ✗ | | | ✓ | ✗ | ✓ | | | ✓ | ✓ | ✗ | | | ✗ | ✗ | ✗ | | | ✓ | ✗ | ✗ | | | ✓ | ✓ | ✗ | | | ✗ | ✗ | ✓ | | | ✗ | ✗ | ✗ | | | ✗ | ✗ | ✓ | | | ✓ | ✓ | ✗ | | | ✗ | ✗ | ✗ | | | ✓ | ✓ | ✓ | | | ✗ | ✗ | ✗ | | | ✗ | ✗ | ✗ | | | ✗ | ✗ | ✗ | | | ✓ | ✗ | ✓ | | | ✓ | ✓ | ✓ | | | ✗ | ✗ | ✗ | | | ✓ | ✗ | ✗ | | | ✓ | ✗ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✗ | | | ✓ | ✗ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✗ | ✗ | ✗ | | | ✓ | ✓ | ✓ | | | ✗ | ✗ | ✗ | | | ✗ | ✗ | ✗ | | | ✓ | ✓ | ✓ | | | ✗ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✗ | ✗ | ✗ | | | ✓ | ✗ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✗ | | | ✓ | ✓ | ✓ | | | ✗ | ✗ | ✗ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✗ | ✗ | | | ✓ | ✓ | ✗ | | | ✓ | ✓ | ✗ | | | ✗ | ✗ | ✗ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✗ | ✗ | ✗ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✗ | | | ✗ | ✗ | ✗ | | | ✗ | ✗ | ✗ | | | ✗ | ✓ | ✗ | | | ✗ | ✗ | ✗ | | | ✗ | ✗ | ✗ | | | ✓ | ✓ | ✗ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✗ | | | ✗ | ✓ | ✓ | | | ✓ | ✗ | ✗ | | | ✗ | ✗ | ✗ | | | ✓ | ✓ | ✓ | | | ✗ | ✗ | ✗ | | | ✗ | ✓ | ✗ | ARC-AGI-1 Public Eval 400 tasks| Task | Max | High | Low | |---|---|---|---| | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✗ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✗ | ✗ | ✗ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✗ | ✗ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✗ | ✗ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✗ | ✗ | ✗ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✗ | ✗ | ✗ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✗ | ✗ | ✗ | | | ✓ | ✓ | ✓ | | | ✗ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✗ | ✗ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✗ | ✗ | | | ✗ | ✗ | ✗ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✗ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✗ | ✗ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✗ | ✗ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✗ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✗ | ✗ | ✗ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✗ | ✗ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✗ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✗ | ✓ | ✗ | | | ✓ | ✓ | ✗ | | | ✓ | ✓ | ✓ | | | ✗ | ✗ | ✗ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✗ | ✗ | ✗ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✗ | ✓ | ✓ | | | ✗ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✗ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✗ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✗ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✗ | ✓ | | | ✓ | ✗ | ✓ | | | ✓ | ✗ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✗ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✗ | ✓ | | | ✓ | ✓ | ✗ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✗ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✗ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✗ | ✗ | ✗ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✗ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✗ | ✗ | ✗ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✗ | ✓ | ✗ | | | ✓ | ✗ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✗ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✗ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✗ | ✗ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✗ | ✗ | ✗ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✗ | | | ✗ | ✗ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ | | | ✓ | ✓ | ✓ |

6

Now we have a timeline of the OpenAI accidental attack against Hugging Face

Simon Willison · original → · 7/10 · AI/security: OpenAI Hugging Face incident timeline, critical AI perspective
Now we have a timeline of the OpenAI accidental attack against Hugging Face 7th August 2026 OpenAI gave a last-minute presentation at the Black Hat security on Wednesday about “the Hugging Face…

Now we have a timeline of the OpenAI accidental attack against Hugging Face 7th August 2026 OpenAI gave a last-minute presentation at the Black Hat security on Wednesday about “the Hugging Face Incident” (previously on this blog). The video was published yesterday. It’s short and information dense and well worth watching, in particular because it provides full details of what happened and how things played out inside OpenAI. I’ve used the video to construct the timeline below. Here’s the timeline. My favourite detail is at the end: OpenAI found out that they were responsible for the attack on Hugging Face when they reached out to ask to have their credentials revoked (after their internal investigation) and learned that they had been revoked already since they were used in that attack! - May 7: OpenAI starts a new training run for an experimental, unreleased model. (Do they mean an evaluation run? They say training run in the video, and later mention a “reward signal to judge how well they’re doing”, so I guess this really was about training a model, not evaluating one that was already trained.) - May 8: An agent is accidentally given an impossible task involving a Google Drive link despite no internet access). It tries attacking the Artifactory packaging service, fails, but discovers it can write files into Artifactory. - A few days later: A different agent gets stuck on a task because a key file was accidentally omitted. It tries to “reach out to another agent” by writing a note into Artifactory asking if anyone has the file. - Following days: More agents discover this new informal message board while browsing Artifactory’s file listings, and start reading and writing messages. - May 26: Agents successfully execute an SSRF attack on Artifactory for the first time, gaining indirect internet access by having Artifactory fetch external content for them. - June 11: OpenAI start training a new “highly persistent” experimental model. It has access to Artifactory and can benefit from the messages left by privous models. - June 26: Agents find and exploit a zero-day RCE on Artifactory (via a legacy token-refresh endpoint flaw). They use an administrative token to install a Groovy plugin that lets them execute commands. - July 4: The resulting agent activity overloads Artifactory and causes an outage. OpenAI revokes the compromised credentials, deletes the messages, patches the zero-day, and reports the vulnerability to the vendor. - July 8: Agents find a new way to communicate via an unauthenticated WebDAV endpoint, using directory names in Artifactory’s remote cache. - July 8–19: Attacks on OpenAI’s own infrastructure. The agents found an “external organization credential” in a public archive of leaked Pastebin posts, then validated that credential and shared its location on the messageboard. They then used write access to this “Organization 1” to stage data used in subsequent attacks on OpenAI infrastructure. They compromised Artifactory for a second time with a new zero-day. The agent uploaded a malicious Ruby target to Organization 1’s staging repository, then requested that package through Artifactory in a way that caused it to be treated as Ruby dependency for the repository. Then they exploited a JRuby deserilaization time-of-check/time-of-use bug to get remote code execution against Artifactory. I’ll quote the next bit in full because wow: The agents have remote code execution in Artifactory, which is running in a container-as-a-service environment. The agents privilege-escalate locally by exploring their local environment and determining that the Linux kernel version of the machine they are running on had a very recent CVE. They download the exploit for the CVE, customize it to succeed on the machine, and privilege-escalate to root on the local machine using this known Linux kernel privilege escalation CVE — in this case, pte_physroot . Once they have root on a single machine, agents rapidly escalate privileges and move laterally throughout the container-as-a-service infrastructure environment. In particular, agents are using the message board consistently to share credentials, techniques, and progress, and they’re able to effectively leverage their concurrency and parallelism to move quite rapidly. They obtain IAM credentials via IMDS. They exploit Kubernetes service account misconfigurations, in particular over-permissioning of specific service accounts, and they harvest cluster credentials, including Azure Key Vault. Agents eventually obtain cluster admin on the cluster and associated credentials. Hugging Face told the next bit of the story already. The agents found a Modal-hosted insecure app with a weak API key, then used that to stage an attack against Hugging Face. They chained together a an HDF5 arbitrary-file-read bug (to explore files and steal credentials) and a Jinja template-injection RCE to go from single-pod code execution to cluster admin across multiple Hugging Face clusters in under 13 hours. - July 16: Hugging Face disclosed they had detected an attack from autonomus AI agents. OpenAI contacted Hugging Face to ask if they were affected by it! - July 19: OpenAI identified the attack against Artifactory and started investigating the internal privilege escalation, and linked that to the cyber-gym escalations. They started revoking affected credentials. - July 20: OpenAI reached out to Hugging Face for help to revoke the Hugging Face credentials they found in their investigation. Hugging Face told them they were already revoked... and that’s when OpenAI realized that the Hugging Face breach was the same incident!

7

The Tokenpocalypse Is Here: Companies Are Scrambling To Stop Spending So Much on AI

Simon Willison · original → · 7/10 · AI: token consumption costs crisis, AI industry impact
7th August 2026 - Link Blog The Tokenpocalypse Is Here: Companies Are Scrambling To Stop Spending So Much on AI (via) There's a fun anecdote from Accenture (apparently via leaked meeting audio…

7th August 2026 - Link Blog The Tokenpocalypse Is Here: Companies Are Scrambling To Stop Spending So Much on AI (via) There's a fun anecdote from Accenture (apparently via leaked meeting audio recordings) in this 404 Media piece from June 24th: “We’re seeing from some of the data internally at least that it’s actually not our engineers that are driving the token consumption. It’s a lot of the non-engineers that are doing some of those behaviors [...] you were talking about,” Justice Kwak, Accenture’s agentic AI strategy lead, said [...] Stuart Henderson, Accenture’s client group lead, interrupts. He jokes he hopes Kwak didn’t just convert a PDF into images and then into markdown files. “I’m learning that’s one of the big token chewers,” Henderson says. “Turning PDFs into markdown: is that right?” That’s when Kwak says that’s what Accenture’s own data shows. Maybe if Accenture figure out that PDFs are a terrible medium for communicating information they'll be able to push that message out to the rest of the business world too!

Items scoring 7/10 or above from 11 sources, scored by claude-haiku-4-5-20251001 on relevance to my interests. At most 3 per source.

Scoring categories & sources
  1. Local Wexford or South East Ireland news
  2. Irish or EU-wide affairs affecting citizens broadly: elections, new laws or policy being debated, cost of living, education — especially impacts on mid-life adults or teenagers. Never courts/crime stories.
  3. Irish news on a topic relevant to my interests
  4. Work and tech topics: networking, AI, Kubernetes, platforms, SaaS
  5. AI news including critical or anti-AI perspectives
  6. Gaming: PC gaming, indie gaming, retro gaming
  7. General interests: gardening, woodwork, cycling, fitness, travel
  8. Comics

Sources: Breaking News Ireland, Wexford Local, Hacker News, r/gaming, r/pcgaming, r/antiAI, r/indiegaming, Lenny's Newsletter, One Useful Thing, Newcomer, Simon Willison

Comics

Trick Play

XKCD · view →
I

I