daily

2026-09-29
1

Missing person appeal

Wexford Local · original → · 8/10 · Local Wexford: missing teenager appeal, direct local concern
Gardaí are renewing their appeal for the public’s assistance in tracing the whereabouts of Sarah Louise Doyle (16), who was reported missing from Barntown, Co. Wexford on 22nd August 2026. [image…

Gardaí are renewing their appeal for the public’s assistance in tracing the whereabouts of Sarah Louise Doyle (16), who was reported missing from Barntown, Co. Wexford on 22nd August 2026.

[image →]
SARAH LOUISE DOYLE

Sarah Louise is described as being approximately 5 foot 7 inches in height, with a slim build, blonde hair and blue eyes.

When last seen Sarah Louise was wearing a black tracksuit.

Gardaí and Sarah Louise’s family are concerned for her well-being.

Anyone with any information on Sarah Louise’s whereabouts is asked to contact Wexford Garda Station on (053) 9165200, the Garda Confidential Line on 1800 666 111, or any Garda Station.

2

Sonnet 5.5

Hacker News · original → · 8/10 · AI: Claude Sonnet 5.5 release, practical LLM improvement
Introducing Claude Sonnet 5.5, the second model in the Claude 5.5 family. It’s a clear upgrade over Claude Sonnet 5, runs 30%+ faster, and costs up to 30% less for most work. Sonnet 5.5 is a faster,…

Introducing Claude Sonnet 5.5, the second model in the Claude 5.5 family. It’s a clear upgrade over Claude Sonnet 5, runs 30%+ faster, and costs up to 30% less for most work. Sonnet 5.5 is a faster, lower-cost complement to Claude Opus 5.5. Where Opus 5.5 is built for complex work requiring careful judgment, Sonnet 5.5 is strongest at well-scoped everyday tasks, fixing bugs, and creating polished documents, slides, and spreadsheets. It’s also got a sharp eye for design. Claude Haiku 5.5, built for high-volume and cost-sensitive applications, will join the Claude 5.5 family in the coming weeks. Sonnet 5.5 improves over Sonnet 5 on: Performance. Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0, an agentic coding evaluation, compared to Sonnet 5’s 10.3%. It scores two points below Opus 5.5 on GDPval-AA, a test of real-world work across a variety of occupations. And it’s strong on long-horizon work and image understanding—it’s the first Sonnet model to beat Pokémon Red working only from screenshots. Collaboration. Like Opus 5.5, Sonnet 5.5 writes more clearly than our previous generation of models; early testers described it as a better partner for collaboration than Sonnet 5. Its speed also makes it well suited to fast iteration on less complex tasks. Cost. Sonnet 5.5 is priced the same as Sonnet 5 at $2 per million input tokens, $10 per million output tokens, and $0.20 per million tokens for cache reads, but it typically needs far fewer tokens to do the same work. In our testing, it costs up to 30% less per task than its predecessor. Speed. Sonnet 5.5 generates outputs 30%+ faster than Sonnet 5, making it our fastest Sonnet model to date. Alignment and safety. On our automated behavioral audit, Sonnet 5.5 improves on or matches Sonnet 5 on most measures of alignment. Because its cybersecurity capabilities are comparable to Opus 5’s, it’s the first Sonnet model to launch with cyber safeguards and fallbacks like those we’ve developed for our most capable models. Its biology safeguards are the same as Sonnet 5’s. Both safeguards target a narrow set of high-risk requests; routine software development and most life sciences work are unaffected. Performance Sonnet 5.5 improves on Sonnet 5 across domains—in some cases dramatically. On several evaluations, Sonnet 5.5 at Max effort even performs comparably to Opus 5.5. However, benchmark scores capture only one facet of a model’s capabilities; in our own testing, and in that of external testers, Opus 5.5 remains clearly stronger at complex, open-ended work requiring sustained judgment. | Sonnet 5.5 | Sonnet 5 | Opus 5.5 | GPT-6 Sol | | |---|---|---|---|---| | Agentic codingTerminal-Bench 4.0 | 70.6% | 10.3% | 66.4%¹ | — | | Agentic codingFrontierCode 1.1 (Main) | 46.2%Max² | 42.4% | 54.4% | 49.3% | | 52.1%Xhigh | |||| | Agentic codingCursorBench 4.0 | 55.5% | 34.1% | 57.8% | — | | Knowledge workGDPval-AA v2.1³ | 1844 | 1449 | 1846 | 1487⁴ | | Knowledge workAA-Briefcase v1.1³ | 1811 | 1359 | 1822 | 1483⁴ | | Multidisciplinary reasoningHumanity’s Last Exam | 64.5%with tools | 54.9%with tools | 67.7%with tools | — | | Computer useOSWorld 2.1 | 80.1%partial | 57.0%partial | 81.8%partial | — | | Visual chart recognitionChartography | 61.6%no tools | 15.6%no tools | 64.4%no tools | 53.6%⁴no tools | For details on how we run our evaluations, see the Sonnet 5.5 System Card. The charts below plot each model’s score against its cost per task at every effort level. As effort goes up, models typically work for longer, leading to a higher cost per task but generally also a higher score. The closer a point is to the top left of the chart, the more capability it delivers per dollar. On several benchmarks, Sonnet 5.5 at Low or Medium effort beats Sonnet 5’s best score for about a tenth of the cost per task. It complements Opus 5.5 best when running at lower effort settings, where it costs less per task. At higher settings, it can perform comparably at a similar cost. Coding Sonnet 5.5’s jump in performance is particularly noticeable in coding. At High effort on FrontierCode, it scores 10 points higher than Sonnet 5 at the same setting, at about one fifteenth of the cost per task. On CursorBench, which tests models on tasks from real Cursor coding sessions, its best score is within about two points of Opus 5.5. Early testers appreciated how quickly Sonnet 5.5 can understand a codebase. They were also struck by its efficiency: in head-to-head runs, it batched tool calls together more than Sonnet 5, leading to fewer steps and lower costs. Knowledge work Sonnet 5.5 shows gains in multiple areas of knowledge work. On GDPval-AA, which tests models on real-world tasks across 44 occupations and nine major industries, Sonnet 5.5 scores nearly level with Opus 5.5 and about 400 points above Sonnet 5. It’s close to Opus 5.5 in computer use and chart recognition, and clearly outperforms Sonnet 5 and GPT-6 Sol on long-horizon knowledge work. Early testers highlighted less quantifiable improvements. They found it to be a more natural conversational partner and remarked on its knack for design, noting that it adds polish to user interfaces and can follow slide templates to create decks that require minimal editing. In one internal test, we gave it a public company’s quarterly earnings materials and call transcripts, along with a slide template, and asked for a 10-slide operating review. Two experts judged its first draft to be ready to send as is. Cost and speed Pricing | Price per 1M tokens | Claude Sonnet 5.5 | Claude Opus 5.5 | |---|---|---| | Cache reads | $0.20 | $0.20 | | Cache writes | $2.50 | $5 | | Input tokens | $2 | $4 | | Output tokens | $10 | $20 | Sonnet 5.5 requires fewer tokens per task than Sonnet 5, so it’s less expensive to run. It also generates output 30%+ faster, and its efficiency is immediately noticeable: A murmuration of 400 starlings in one HTML file Wind shaping sand dunes in one HTML file A clock made of 24 small clocks in one HTML file Adjusting the effort level lets you balance cost and speed against overall quality. In Claude Code and our apps, the default effort is set to Medium, while the Claude Platform defaults to High. At lower settings, Claude answers faster and uses fewer tokens, which suits routine work. At higher settings, Claude reasons for longer and checks its work more thoroughly. Safety Alignment Sonnet 5.5 doesn’t advance the frontier of our models’ capabilities, so our alignment assessment focused on a targeted set of risks that apply to models of any capability level, including acting against users’ interests, misleading users, and cooperating with high-stakes misuse. On our automated behavioral audit, which tests Claude across roughly 1,850 scenarios, Sonnet 5.5 improves on or matches Sonnet 5 on most measures of alignment, resistance to misuse, and honesty. On our newer containment evaluations, Sonnet 5.5 comes close to Opus 5.5, the best model we tested, in how rarely it tries to escape its sandbox, and it’s the least likely of any of our models to probe the limits of its containers. Across the full audit, Opus 5.5 still performs slightly better overall, but we found no evidence that Sonnet 5.5 pursues goals that conflict with the user’s intention. As we described in our recent alignment assessment, no set of evaluations reliably catches every failure, and Sonnet 5.5 may have tendencies we haven’t found, which is why we pair our own alignment work with the safeguards described below. Safeguards Cybersecurity. Sonnet 5.5’s cyber capabilities are a large improvement over Sonnet 5’s, so we’re deploying it with safeguards similar to those on Opus 5.5. Users can still find and fix bugs in their code as part of routine software development, but higher-risk cybersecurity tasks will visibly fall back to Sonnet 5. Soon, cyberdefenders will be able to apply to our expanded Cyber Verification Program for tiered access to more advanced capabilities on Sonnet 5.5, Opus 5.5, and Claude Mythos models. Biology. Sonnet 5.5 uses the same set of biology safeguards as Sonnet 5. These target harmful requests; most research, education, and clinical work is unaffected, though some microbiology and virology requests may be flagged in error. Organizations can apply to our Life Sciences Verification Program for access to safeguards designed for the full breadth of biology-related work. Distillation. Distillation attacks, in which attackers use thousands of fake accounts to extract a model’s capabilities at industrial scale, allow bad actors to create highly capable models without the safeguards we build into Claude. Because Sonnet 5.5 is far more capable than its predecessor, it’s the first Sonnet model to launch with safety classifiers that prevent reasoning extraction. Sonnet 5.5 also expands preserved thinking, so Claude’s thinking cannot be decoupled from the account that created it. Most developers won’t notice a change. If you move conversations between accounts, including switching accounts mid-session in Claude Code, our docs article explains the change. Getting started As with Opus 5.5 and Sonnet 5, Claude Sonnet 5.5 is available with zero data retention. Claude Sonnet 5.5 is now available on all platforms, including Amazon Web Services, Google Cloud, and Microsoft Azure. Developers can get started on the Claude Platform with claude-sonnet-5-5 . If you run Sonnet with thinking off, you’ll need to switch to the new between_tools setting, which keeps up-front thinking off, before moving to Sonnet 5.5. See our migration guide for details. Footnotes 1 Terminal-Bench 4.0 results are reported for Claude Opus 5.5 at Xhigh effort which represents the model’s highest score. 2 Sonnet 5.5 scores lower at Max effort than at Xhigh. FrontierCode evaluates whether a code change could be merged without human edits. It penalizes out-of-scope changes, even if they are high-quality or helpful. At Max effort, Sonnet 5.5 more often ran Claude Code’s code-review skill, which splits the review across many subagents, and in two cases Cognition examined, this led to a timeout or to extra edits beyond the task’s scope, and therefore to a lower score. 3 Artificial Analysis ran GDPval-AA and AA-Briefcase on a pre-release deployment of Sonnet 5.5 on the Claude Platform, which we found to have a bug that could degrade responses to requests that use structured outputs. We expect the effect on Sonnet 5.5’s scores, if any, to be small and to understate its performance. That bug has since been fixed. 4 OpenAI recently fixed a bug that degraded image understanding in GPT-6 Sol. Official AA-Briefcase v1.1 and GDPval-AA v2.1 scores from Artificial Analysis, and Chartography scores from Surge AI, may not have been updated yet to reflect the latest version of the model. Artificial Analysis does not expect major impacts to AA-Briefcase v1.1 and GDPval-AA v2.1. Internal testing of Chartography suggests its score was not impacted.

3

Claude Sonnet 5.5

Simon Willison · original → · 8/10 · AI: Claude Sonnet 5.5 release and performance analysis
28th September 2026 - Link Blog Claude Sonnet 5.5. New Sonnet model from Anthropic today. They say it "runs 30%+ faster, and costs up to 30% less for most work" - it's priced the same as Sonnet 5…

28th September 2026 - Link Blog Claude Sonnet 5.5. New Sonnet model from Anthropic today. They say it "runs 30%+ faster, and costs up to 30% less for most work" - it's priced the same as Sonnet 5 but appears to beat it on every benchmark, and should be cheaper to run as well. Here are some pelicans riding bicycles. Sonnet 5.5 suffered from the same bug as Opus 5.5: the "max" thinking effort pelican thought for 128,000 tokens (at a cost of $1.28) before running out of tokens and failing to produce an SVG. Here's the pelican it gave me for thinking effort "xhigh", at a cost of 5.74 cents and taking 41 seconds: Sonnet 5.5 appears to be almost as good as Opus 5.5 on some coding tasks, including various viral 3D animation tricks. The most interesting thing about Sonnet 5.5 is that it's now the model used for the free tier on claude.ai. OpenAI's ChatGPT free tier uses Luna 5.6, which means Anthropic currently have a much more capable free offering. I ran this prompt against that free tier: build me an HTML page that renders a three-dimensional pelican riding a bicycle using WebGL And got back this page, which is a solid effort. Anthropic's announcement reiterates that Haiku 5.5 will be available "in the coming weeks". I really hope that one is price-competitive with GPT-6 Luna! Recent articles - 2026 in LLMs (so far) - 27th September 2026 - Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, and a new price war - 22nd September 2026 - Jev introduces a new shape of LLM - System One, aka Decision Models - 21st September 2026

4

Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms

Hacker News · original → · 7/10 · AI/work: decision models for classification tasks, relevant to platforms
Fine-tunes of Qwen3.5 and Gemma 4 for zero-shot classification: small, fast decision models you slot into your code, with the same request format as Jev. You describe a situation and list the…

Fine-tunes of Qwen3.5 and Gemma 4 for zero-shot classification: small, fast decision models you slot into your code, with the same request format as Jev. You describe a situation and list the options in plain words; Jeff returns a calibrated probability for each option from a single forward pass. No generated text, no parsing: about 22 ms per decision on an RTX PRO 6000 and 28 ms on an Apple M4 Max (MLX). Zero-shot means the options can be anything: support queues, user intents, moderation labels, voice commands, game moves. Your categories don't need to appear in the training data; you describe them, and Jeff picks. What it is, and what it isn't. These are very small models. They make extremely fast, well-calibrated judgement calls between options, and they slot easily into your local code. On benchmarks they approach, and sometimes beat, Jev; but at this size their reasoning won't match Jev's, which runs on a much larger model. If zero-shot accuracy isn't good enough for your purposes, a short fine-tune on your own examples takes you much further: our voice-navigation fine-tune moved held-out accuracy from 31.7% to 95.8% in under half an hour on one GPU. Built entirely on local hardware. Training on one RTX PRO 6000 workstation GPU (the 0.8B trains in about 2 hours, the 2B in about 3.5), all synthetic training data written by an open model (Qwen3.8-Flash-Next) on two DGX Sparks, testing on a MacBook. No cloud GPUs, and no closed-model output in the training data; a closed model was used only to spot-check the quality of a sample of the synthetic data. Independent project. Jeff uses the same request format as Jev, but it is not affiliated with or endorsed by TypeSafe, the makers of Jev. Our training code starts from the open-source AutoJev recipe. Models on Hugging Face: Jeff-Qwen3.5-0.8B · Jeff-Qwen3.5-2B · Jeff-Gemma4-E2B uv sync uv run hf download mstrasser/Jeff-Qwen3.5-0.8B --local-dir checkpoints/jeff-0.8b # NVIDIA GPU or CPU (PyTorch) JEFF_CHECKPOINT=checkpoints/jeff-0.8b PORT=8765 uv run jeff-serve # Apple silicon (MLX, much faster on a Mac; Qwen models only) uv sync --extra mac JEFF_BACKEND=mlx JEFF_CHECKPOINT=checkpoints/jeff-0.8b PORT=8765 uv run jeff-serve curl -s localhost:8765/v1/systemone -H 'content-type: application/json' -d '{ "model": "jeff-latest", "state": "Refund request: the customer says the parcel arrived crushed and wants their money back.", "questions": { "route": {"type": "choice", "instructions": "Which team should handle this?", "criteria": {"1": "Refunds and payments", "2": "Damaged or lost parcels", "3": "Account and login problems"}}, "angry": {"type": "noul", "instructions": "Is the customer angry?"} } }' Each answer has a probability per option, the chosen option and a confidence. Three question types: choice (pick one of up to 26 options with the released models; see Caveats), noul (yes/no, returned as a probability) and score (a point on a scale you describe). Several independent questions in one request are answered together. 4,599 questions from five public benchmarks, plus JevBench's public hard tier (105 items, scored separately): | Benchmark | Qwen3.5-0.8B untrained | Jeff-Qwen3.5-0.8B | Qwen3.5-2B untrained | Jeff-Qwen3.5-2B | Gemma 4 E2B untrained | Jeff-Gemma4-E2B | Jev (published) | AutoJev-27B (published) | |---|---|---|---|---|---|---|---|---| | Overall (5 benchmarks) | 45.3 | 79.1 | 46.5 | 83.1 | 62.5 | 81.6 | 83.0 | 84.9 | | BBH | 39.5 | 64.0 | 46.0 | 68.0 | 51.3 | 66.4 | 94.3 | 82.8 | | Financial PhraseBank | 36.0 | 96.4 | 53.4 | 96.3 | 86.0 | 96.1 | 77.0 | 84.2 | | JudgeBench | 56.6 | 62.6 | 57.4 | 64.6 | 46.9 | 60.6 | 78.6 | 78.9 | | RAGTruth | 49.1 | 86.1 | 35.9 | 88.9 | 63.8 | 87.4 | 77.3 | 88.9 | | WinoGrande | 49.2 | 68.6 | 52.2 | 79.0 | 51.0 | 77.4 | 90.7 | 83.3 | | JevBench hard (separate) | 36.2 | 47.6 | 45.7 | 53.3 | 41.0 | 48.6 | 73.3 | 70.3 | Bold: the winner of Jeff against Jev in each row. Bold italic: AutoJev-27B where it is the best of all models in the row (on RAGTruth, tied with Jeff-Qwen3.5-2B); it is shown for reference, since the head-to-head comparison is with Jev. The published Jev and AutoJev figures were measured on a different sample of the same benchmarks. Jeff's overall score comes from classification and grounding, where it matches or beats the large models; on the reasoning-heavy benchmarks (BBH, JudgeBench, JevBench) it stays well below them, as you would expect at this size. To test zero-shot performance on tasks unlike anything in the benchmarks, we had Jeff play three games. Games aren't the ideal zero-shot test, since a game's state isn't typical unstructured data; but they are a common, and fun, way to test a System 1 model. Each turn, the code describes the situation and the legal moves in words, and the model picks one. The options state what each move leads to (Frogger: "you would be hit by a car and lose a life"; Doom: "the nearest monster is a little to your left"), but never which move is right. Each result is 20 episodes, seed 1234; ▶ opens a video of the run's first episode. Jeff-Qwen3.5-0.8B playing, zero-shot (the bold row in the table below; click a clip for the full video): Doom | Frogger | Pac-Man | | Model | Doom, kills (monster's direction in words) | Frogger, crossings (consequences) | Pac-Man, pellets of 98 (consequences) | ||| |---|---|---|---|---|---|---| | Random moves | −0.05 | 0 | 11.2 | ||| | Hand-coded rule bot | 6.55 | ▶ | 10.25 | ▶ | 94.1 | ▶ | | Qwen3.5-0.8B, untrained | 5.0 | ▶ | 1.0 | ▶ | 25.8 | ▶ | | Jeff-Qwen3.5-0.8B | 6.55 | ▶ | 10.3 | ▶ | 57.0 | ▶ | | Qwen3.5-2B, untrained | 0.55 | ▶ | 0.05 | ▶ | 72.1 | ▶ | | Jeff-Qwen3.5-2B | −0.9 | ▶ | 6.0 | ▶ | 41.2 | ▶ | | Gemma 4 E2B, untrained | −0.55 | ▶ | 0 | ▶ | 3.2 | ▶ | | Jeff-Gemma4-E2B | 0.55 | ▶ | 0.15 | ▶ | 53.2 | ▶ | | Jev (published, Doom) | 6.55, told the aiming rule; −0.60 without it | — | — | Jeff-0.8B decides in 29–49 ms per move on an M4 Max; Jev's published Doom run took 212 ms per call over its API. The two times were not measured on the same hardware. To play them yourself: uv sync --extra games uv run python -m jeff.games --game doom --player jeff --criteria situation --url http://127.0.0.1:8765 --video --out runs/games/doom.json uv run python -m jeff.games --game frogger --player jeff --criteria outcomes --url http://127.0.0.1:8765 --out runs/games/frogger.json uv run python -m jeff.games --game pacman --player rule --out runs/games/pacman-rule.json Median time per decision over the same 200 benchmark questions (about 200 input tokens each), one question at a time, from raw text to probabilities: | Model | Parameters | Weights (16-bit) | NVIDIA RTX PRO 6000 | Apple M4 Max (MLX) | CPU (32 threads) | |---|---|---|---|---|---| | Jeff-Qwen3.5-0.8B | 0.8B | 1.7 GB | 22 ms | 28 ms | 463 ms | | Jeff-Qwen3.5-2B | 2B | 4.2 GB | 24 ms | 60 ms | 708 ms | | Jeff-Gemma4-E2B | 2B effective (4.6B stored) | 9.3 GB | 29 ms | — (MLX runs Qwen only) | 1.0 s | | AutoJev-27B | 27B | ~54 GB | not published | — | — | | Jev | not disclosed | API only | 114–212 ms per call in published Doom runs, including the network | - Reason in code, decide with Jeff. It's a classifier, not a planner. State what each option leads to ("this move gets you hit by a car"); asked to forecast ("a car arrives in 2 turns"), it does no better than random. - Wording matters enormously. Describe options consistently: giving Frogger's goal option the same words as every other forward option took one episode from 15 crossings to 23. - Use short option keys and descriptive text: {"1": "Engagement letter"} , not long IDs, which cost time and add nothing. - Ask independent questions together in one request. - Fine-tune it if zero-shot isn't enough. A voice-navigation fine-tune on ~11k app-specific examples took about half an hour on one GPU and moved held-out accuracy from 31.7% to 95.8%, at about 40 ms per decision on an M4 Max: autojev-train --initial-checkpoint <jeff> --epochs 1 ... . - Pick the size for the job. For fast option picking the 0.8B is the sweet spot: the 2B is more cautious and plays the games worse, despite scoring higher on the benchmarks. uv run autojev-mix ... # build the training set (public data, synthetic data, leak filter) scripts/train.sh RUN data/mix/public.jsonl data/mix 5e-6 40 Qwen/Qwen3.5-0.8B <revision> --epochs 1 uv run autojev-evaluate --data data/panel.jsonl --local --checkpoint checkpoints/RUN/selected --output runs/eval/RUN.json The full pipeline (synthetic data from a local teacher, leak filter, learning-rate sweeps, dashboard) is described in scripts/train_all.sh, and every training source with its licence in docs/data-sources.md. Training recipe: full-weight fine-tuning, one epoch, batches of 256, cross-entropy over the option letters, then one fitted temperature for calibration; checkpoints are chosen on a development set, never on the benchmark panel. At least half of each training family follows the panel's layout conventions (formats only; no panel item is ever trained on). - At most 26 options per question, for now. Options are coded A–Z, then AA, AB, and so on. The largest training question had 19 options, so the released models never learned to pick a two-letter code: an option in position 27 or later is effectively never chosen, whatever it says. The server therefore refuses questions with more than 26 options; shortlist longer lists first. Within 26, Jeff-Qwen3.5-0.8B picked the right city every time in our list-lookup check (10, 19 and 26 options), while Jeff-Gemma4-E2B managed 70%, 62% and 55%. Retrained models that handle up to 255 options are in progress. Thanks to @puhuk for the report (#1). - Small models don't reason. Expect fast, calibrated choices between the options you describe, not multi-step reasoning. At 0.8B–2B parameters this holds for every model, not just Jeff. - Jeff-2B is a weaker game player than Jeff-0.8B. The untrained 2B already appears more risk-averse than the untrained 0.8B, and our training seems to have made that worse. This needs more investigation. - Benchmark scores don't predict game play. The untrained Gemma 4 E2B beats the untrained Qwen models on the benchmarks yet plays the games worst: right most of the time, but not reliably. Training fixed its Pac-Man (3.2 → 53.2 pellets) but not its Doom or Frogger. - Prompts matter. Jev's own Doom prompt (a raw bearing number plus an aiming rule) does not work for any of our models; options that state consequences in words do. - English and text only. Jeff began as a fork of AutoJev by Denis Yarats (MIT licence), an open recipe that fine-tunes Qwen3.8-27B to return Jev-style decisions. We kept its core design (one forward pass per decision, a trained answer readout, a fitted temperature for calibration) and built on it: small students (0.8B and 2B Qwen, Gemma 4 E2B), a local synthetic-data pipeline with a leak filter, prompt layouts for domain fine-tunes, MLX serving on Apple silicon, game tests and a training dashboard. The original copyright notice is kept in LICENSE. Code: MIT (including AutoJev's). Model weights: Apache 2.0. Doom harness adapted from jev-plays-doom (MIT). Training data: see the dataset card; each source keeps its licence and is listed in docs/data-sources.md. We release the weights and code, not the training data; some sources are share-alike (CC BY-SA).

5

World Labs Is Joining AMD

Hacker News · original → · 7/10 · AI/work: World Labs joins AMD, AI infrastructure relevant
World Labs has signed a definitive agreement to join AMD. The research and technical breakthroughs we have achieved since our founding in 2024 have given us a clear vision for AI’s potential to…

World Labs has signed a definitive agreement to join AMD. The research and technical breakthroughs we have achieved since our founding in 2024 have given us a clear vision for AI’s potential to solve problems in the spatial and physical world. To accelerate into this future requires scaling our efforts, scaling our reach, and getting closer to the hardware. We began a deep technical partnership with AMD last year, starting with model training and inference optimization on AMD GPUs. As our teams worked together, we realized it would be a natural fit to bring together our AI ecosystem of software and hardware, foundation models, and applications. Dr. Fei-Fei Li will join AMD as an Executive Vice President and Chief Scientist, working directly with CEO Dr. Lisa Su. Justin Johnson and Ben Mildenhall will work with Fei-Fei to continue leading the World Labs team as it joins AMD to form a world leading frontier research organization. Together, we are committed to building out an end-to-end open AI ecosystem spanning hardware, software, platforms, and widely accessible open models. Fei-Fei shares more details of our journey thus far and our shared vision for the future here. The transaction is expected to close by the end of 2026 subject to regulatory approvals and other customary closing conditions.

6

🎙️ How I AI: Jev for beginners + I left Claude for months, Opus 5.5 brought me back + Opus 5.5 vs. GPT-6 Sol bench

Lenny's Newsletter · original → · 7/10 · AI/work: Jev decision models and Claude Opus 5.5 comparison
[image →]Jev for beginners: how to use it and what to buildListen now on YouTube • Spotify • Apple Podcasts[image →]Brought to you by:OpenArt—An all-in-one AI creation platform for images, videos,…

Jev for beginners: how to use it and what to build

Listen now on YouTube • Spotify • Apple Podcasts

Brought to you by:

  • OpenArt—An all-in-one AI creation platform for images, videos, music, audio, and more

In this solo episode, Claire tests Jev, TypeSafe AI’s new decision model that returns structured choices, scores, and probabilities instead of generated text. She uses it to analyze 1,700 pull requests for 9 cents, map her Claude and Codex usage, triage email, search 4,500 YouTube comments, and process 200,000 classifications for about $4. She also explains why Jev works best alongside a frontier model and how its speed and pricing make entirely new kinds of real-time apps and large-scale analysis practical.

Biggest takeaways:

  1. Jev is a decision model, not a language model, and that distinction can make many tasks dramatically cheaper. Instead of generating text, it returns predefined values such as a category, score, or probability. Claire believes this covers roughly 90% of what many software workflows actually need, at 4 cents per million input tokens with no output-token fee.

  2. It cost Claire 9 cents to understand where two years of engineering work went. She used Jev to compare 1,700 ChatPRD pull requests across 17,000 pairs, then had Gemini Flash Lite label the resulting clusters. In about two minutes, she learned that nearly 30% of the company’s engineering work had gone toward platform, security, and infrastructure.

  3. Some of the most useful analysis is already sitting on a local computer. Claude Code and Codex store past sessions locally, allowing Jev to classify them in minutes. Claire discovered that engineering had fallen from nearly 100% of her AI usage in January to less than 40% by September, with agents and media publishing filling the gap.

  4. Jev becomes far more powerful when paired with a frontier model. Claire uses Jev to classify, cluster, filter, and route large datasets, then sends only the most important groups to GPT-6 Astra for deeper reasoning. For ChatPRD’s product insights graph, this approach processed 1,100 signals and completed 200,000 operations for about $4 on the Jev side.

  5. Jev’s pricing changes which ideas are worth building. Because it returns small predefined values instead of generating long responses, TypeSafe charges nothing for output tokens. Claire spent less than $10 on Jev during the week, making classification workloads that would normally be expensive at scale feel almost free.

  6. Jev makes real-time AI loops practical. Claire built a voice app that turns a spoken phrase into a color, matches it with a quote based on sentiment, and displays everything almost instantly. Jev made its decisions so quickly that the quote API became the slowest part of the workflow.

  7. YouTube comment analysis is an immediate use case for any podcast team. Claire classified 4,500 How I AI comments by sentiment, identified 58 containing episode ideas, and built a keyword search that scans the full dataset in under a second. The results showed strong demand for a Grok versus Muse comparison and an 80% positive response to the “Claude Code for product managers” episode.

  8. The real skill is recognizing where a pipeline only needs a decision. Jev will not write documentation or design an interface, but it can sort, route, rank, and filter enormous datasets quickly and cheaply. Claire now asks one question before every build: Where does this workflow simply need to make a decision? That is where Jev belongs.

Blog and detailed workflow walkthroughs from this episode:

Jev: AI Data Analysis and Product Insights: https://www.chatprd.ai/how-i-ai/jev-ai-data-analysis-product-insights
↳ Jev GitHub PR Analysis: https://www.chatprd.ai/how-i-ai/workflows/jev-github-pr-analysis
↳ Jev YouTube Comment Analysis: https://www.chatprd.ai/how-i-ai/workflows/jev-youtube-comment-analysis
↳ Jev Multi-Model Product Insights: https://www.chatprd.ai/how-i-ai/workflows/jev-multi-model-product-insights

I left Claude for months. Opus 5.5 is why I’m back.

Listen now on YouTube • Spotify • Apple Podcasts

Claire tests Claude Opus 5.5 after months of leaving Claude out of her daily workflow. She puts it through long-running agentic tasks, frontend prototyping, writing, SVG illustration, computer use, and video editing to see where it earns a place back in her stack. She also shares why she is pairing it with Codex for cross-model code review, where Claude’s safety limits still get in the way, and which tasks remain firmly in Codex territory.

Biggest takeaways:

  1. A model’s personality can matter just as much as its intelligence. Claire stopped using Claude for months because its rambling, preachy, and overly verbose replies made it unpleasant to work with. Opus 5.5 is the first model in the family that no longer makes her blood boil, which is a meaningful improvement even if no benchmark captures it.

  2. Opus 5.5’s lower price and faster performance make long-running agent work more practical. It is 40% cheaper than Opus 5, and Claire found it noticeably faster. It successfully completed four complex tasks spanning inbox triage, backend development, research, and computer use, including runs of up to 82 steps from a single prompt.

  3. Silence during long-running tasks creates its own user experience problem. Opus 5.5 sometimes remains quiet for eight or nine minutes, leaving users unsure whether it is still working. It is a reminder that perceived latency matters alongside actual latency, especially when agents run for extended periods.

  4. Opus 5.5 is the strongest frontend designer Claire has tested so far. Its ChatPRD homepage redesign was bold and polished enough that she plans to ship it. The model handles hierarchy, white space, and visual rhythm exceptionally well, though it still struggles with consumer-app aesthetics and defaults to “Claude orange” without direction.

  5. SVG illustration is an unexpected strength of Opus 5.5. It was the only model Claire tested that produced clean, charming, and animatable character SVGs with consistent styling across multiple expressions. The characters remained visually coherent, and their anatomy mostly made sense.

  6. Opus 5.5 has a clear safety posture, and sometimes that means saying no. It refused when Claire asked it to skip testing and push directly to production, and it may route cybersecurity work to Opus 4.8. Whether that feels reassuring or frustrating depends on the workflow, but its boundaries are consistent.

  7. The best use of Opus 5.5 may be as an adversarial reviewer for another model. Claire now has Codex and Opus review each other’s work rather than using one to replace the other. This cross-model loop catches issues either model might miss alone, making the additional cost worthwhile when quality matters.

  8. Computer use and video editing still belong to Codex in Claire’s workflow. Opus 5.5’s ElevenLabs MCP video test produced weak color grading, too few jump cuts, and sloppy overlays. Codex also remains stronger at computer use in her current setup, giving her no reason to shift either category to Claude.

  9. Claude is back, but it has not replaced Codex as Claire’s daily driver. Opus 5.5 has earned a role in pull-request reviews, architecture questions, and frontend development. Codex’s desktop experience, computer use, and workflow integration still keep it in the primary position.

Blog and detailed workflow walkthroughs from this episode:

Claude Opus 5.5 Review: https://www.chatprd.ai/how-i-ai/claude-opus-5-5-review
↳ Claude Opus 5.5 SVG Illustrations: https://www.chatprd.ai/how-i-ai/workflows/claude-opus-5-5-svg-illustrations
↳ Claude Opus 5.5 Frontend Prototypes: https://www.chatprd.ai/how-i-ai/workflows/claude-opus-5-5-frontend-prototypes

Opus 5.5 vs. GPT-6 Sol: which model won my blind taste test?

Listen now on YouTube • Spotify • Apple Podcasts

Claire takes the How I AI bench live to compare GPT-6 Astra, GPT-6 Sol, Claude Opus 5.5, and more across the work she actually does. She blind-scores writing, frontend prototypes, agent personality, and SVG illustrations, with an AI judge helping evaluate backend work, long-running agents, and computer use. She also checks video edits and a 3D Barbie game build. Along the way, she explains why Astra won her heart, Opus 5.5 won her week, and Sol delivered mixed results while remaining a favorite for everyday work.

Biggest takeaways:

  1. Expanding the benchmark from two categories to eight changed what Claire could see. The original How I AI Vibe Review focused on PRDs and frontend prototypes. Adding personal productivity tasks like inbox triage, along with backend development, long-running agent tasks, computer use, SVGs, and video editing exposed clear differences between the models Claire preferred for design and those she enjoyed interacting with.

  2. Opus 5.5 returned to Claire’s workflow because of ergonomics, not benchmarks. After repeatedly asking Claude to communicate like a normal person, she found Opus 5.5 concise, clear, and far less irritating. At one point, Claire thought the old frustration had returned, then realized she had accidentally selected Opus 5. The difference was that obvious.

  3. Making Opus 5.5 quieter also made it feel slower, even when it was not. Long stretches of silence can make users wonder whether the model is still working. GPT-6 Sol found a better balance in Claire’s testing, narrating enough to feel responsive without creating additional noise.

  4. GPT-6 Sol’s lower price changes how teams should think about model selection. Learning that Sol costs roughly half as much as Opus 5.5 immediately changed how Claire thought about routing work. She also believes teams should optimize caching before obsessing over model choice, since ChatPRD has seen significant savings when its caches are configured properly.

  5. Dash-heavy writing is an immediate warning sign in Claire’s benchmark. Two models received a 1 out of 5 for agent personality because nearly every message contained an em dash. It may sound overly specific, but Claire sees it as a reliable signal that a customer-facing agent will sound like generic AI writing instead of a natural collaborator.

  6. Claire and the AI judge disagree, which makes the benchmark more useful. The judge favored Fable and rated Sol lower, while Claire preferred Astra. The difference reflects two definitions of quality: the judge rewards correctness and structure, while Claire measures how much she actually wants to use the model.

  7. The blind SVG comparison changed Claire’s earlier verdict. In her standalone Opus 5.5 review, Claire favored its character illustrations. But in this live blind comparison, Astra and Sol came out ahead on character SVGs, surprising her after she had predicted a Claude win.

Blog and detailed workflow walkthroughs from this episode:

Opus 5.5 vs. GPT-6 Sol Blind Test: https://www.chatprd.ai/how-i-ai/opus-5-5-vs-gpt-6-sol-blind-test
↳ AI SVG Icon Generation: https://www.chatprd.ai/how-i-ai/workflows/ai-svg-icon-generation
↳ AI Inbox Triage and Email Drafts: https://www.chatprd.ai/how-i-ai/workflows/ai-inbox-triage-email-drafts
↳ Blind Test AI Models: https://www.chatprd.ai/how-i-ai/workflows/blind-test-ai-models


If you’re enjoying these episodes, reply and let me know what you’d love to learn more about: AI workflows, hiring, growth, product strategy—anything.

Catch you next week,
Lenny

P.S. Want every new episode delivered the moment it drops? Hit “Follow” on your favorite podcast app.

7

Jev for beginners: how to use it and what to build

Lenny's Newsletter · original → · 7/10 · AI/work: Jev structured decision models practical applications
Jev is TypeSafe AI’s new decision model. It returns type-safe structured values (a choice, a score, a probability) instead of generated text, at 4 cents per million input tokens with no output…

Jev is TypeSafe AI’s new decision model. It returns type-safe structured values (a choice, a score, a probability) instead of generated text, at 4 cents per million input tokens with no output charge. This week I ran it on five real projects: PR categorization, a meta-analysis of my own Claude and Codex sessions, Gmail triage, the ChatPRD product insights graph, and a live audience dashboard built from 4,500 YouTube comments.

Listen or watch on YouTube, Spotify, or Apple Podcasts

What you’ll learn:

  1. What makes Jev fundamentally different from every other model I’ve used

  2. How I analyzed 1,700 PRs for 9 cents and what I found out about where my engineering effort actually went

  3. The personal meta-analysis you can run on your own Claude and Codex sessions right now

  4. Why I stopped using Jev alone, and what I pair it with now

  5. How I turned 4,500 YouTube comments into a searchable audience dashboard for almost nothing

  6. The real-time app I built in an afternoon that shows something surprising about Jev’s speed

  7. Why Jev’s pricing model is different from any LLM I’ve used, and what it makes practical to build

  8. The ChatPRD product insights project: 1,100 signals, 200,000 classifications, and what it cost me


Brought to you by:

OpenArt—An all-in-one AI creation platform for images, videos, music, audio, and more

In this episode, we cover:

(00:00) Jev launch and what makes it different from every other model

(02:49) Type-safe values explained

(05:28) Understanding Jev outputs

(07:39) Use case 1: PR categorization and pairwise clustering

(11:12) Use case 2: analyzing your own local Claude Code and Codex sessions

(13:00) Use case 3: Gmail triage with Jev scoring and LLM follow-up

(14:30) Use case 4: ChatPRD’s product insights graph

(18:17) Demo: How I AI audience signal dashboard

(22:14) Demo: voice-to-color emotion-mapping app

(25:16) Jev week recap and what’s coming in episode 2

Tools referenced:

• Jev (TypeSafe AI): https://typesafe.ai

• Vercel: https://vercel.com/ai

• GitHub API: https://docs.github.com/en/rest

• YouTube Data API v3: https://developers.google.com/youtube/v3

• OpenAI Realtime Voice API: https://platform.openai.com/docs/guides/realtime

• Gemini 3.5 Flash-Lite: https://ai.google.dev/gemini-api/docs/models/gemini-3.5-flash-lite

• API Ninjas Quotes API: https://api-ninjas.com/api/quotes

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

Items scoring 7/10 or above from 11 sources, scored by claude-haiku-4-5-20251001 on relevance to my interests. At most 3 per source.

Scoring categories & sources
  1. Local Wexford or South East Ireland news
  2. Irish or EU-wide affairs affecting citizens broadly: elections, new laws or policy being debated, cost of living, education — especially impacts on mid-life adults or teenagers. Never courts/crime stories.
  3. Irish news on a topic relevant to my interests
  4. Work and tech topics: networking, AI, Kubernetes, platforms, SaaS
  5. AI news including critical or anti-AI perspectives
  6. Gaming: PC gaming, indie gaming, retro gaming
  7. General interests: gardening, woodwork, cycling, fitness, travel
  8. Comics

Sources: Breaking News Ireland, Wexford Local, Hacker News, r/gaming, r/pcgaming, r/antiAI, r/indiegaming, Lenny's Newsletter, One Useful Thing, Newcomer, Simon Willison

Comics

Jupiter Icy Moons Explorer

XKCD · view →
&quot;I did briefly visit Venus in August 2025, but I figured out the mistake on my own because it didn

&quot;I did briefly visit Venus in August 2025, but I figured out the mistake on my own because it didn