daily

2026-08-11
1

Duncannon Beach ‘Do Not Swim’ warning

Wexford Local · original → · 8/10 · Local Wexford: Duncannon Beach water quality warning
[image →]DO NOT SWIM notice has been issued for Duncannon’s Blaue Flag beach today. (File Pic; WexfordLocal.com) By Dan Walsh A “Do Not Swim” notice has been posted at Duncannon Beach today…
[image →]
DO NOT SWIM notice has been issued for Duncannon’s Blaue Flag beach today. (File Pic; WexfordLocal.com)

By Dan Walsh

A “Do Not Swim” notice has been posted at Duncannon Beach today (Monday).

The latest sampling conducted as part of Wexford County Council’s normal monitoring of bathing water quality, showed elevated levels of bacteria at the Beach.

Following consultation with the HSE and EPA, it is necessary to issue “Do Not Swim” warning notices at the above beaches in accordance with the Bathing Water Quality Regulations 2008 and in the interest of public health.

Further samples have been taken and results are expected on Thursday, August 13th at which stage the bathing prohibition notices will be reviewed.  In addition, the Council’s Environmental Technical Team are investigating the matter. 

Wexford County Council advises members of the public visiting the above beach to please abide by the public notices advising against swimming.

2

Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows

Hacker News · original → · 8/10 · AI/work: Meta open agentic model for local workflows
Introducing Muse Glimmer: An Open Agentic Model That Runs on Your Device Today, we're introducing Muse Glimmer, the next model from Meta Superintelligence Labs, and open sourcing the model weights…

Introducing Muse Glimmer: An Open Agentic Model That Runs on Your Device Today, we're introducing Muse Glimmer, the next model from Meta Superintelligence Labs, and open sourcing the model weights under a permissive Apache 2.0 license. Muse Glimmer is a 30-billion-parameter model optimized for always-on local agent workflows. It’s small enough to run on a Mac or PC with a single consumer GPU, enabling use cases that range from local agents and function calling, to local coding, and LLM-as-a-judge evaluation. Muse Glimmer delivers strong performance on key agentic use cases and benchmarks compared with leading models in its size category. Foundation models have achieved remarkable capabilities across reasoning, code generation, and tool use — yet most deployments still depend on cloud infrastructure and network access. Running models locally enables you to use AI anywhere, anytime, with or without an internet connection. This is increasingly viable: the open source community has shown that smaller models, when trained effectively, can approach frontier-level performance on targeted tasks. Muse Glimmer is optimized for these local use cases. Keeping with our long tradition of sharing fundamental AI research, we're releasing Muse Glimmer open weights today on Hugging Face, along with developer documentation to help you start building and running your own agents. Muse Glimmer is built to work with the tools developers already use. Optimized integrations on llama.cpp, MLX, and ExecuTorch will land in the coming days, so you can go from download to working agent in minutes. How We Trained Muse Glimmer An agent that manages your schedule, drafts your messages, organizes your files, and learns how you work needs deep access to personal context. It also needs several capabilities working in concert: long-horizon execution, precise tool calling, multimodal understanding, long-context memory, and instruction following. We designed Muse Glimmer to balance capability against the memory and compute constraints of local hardware. This required a compact architecture, a novel distillation recipe that transfers agentic reasoning from a much larger teacher model, and inference optimizations — including quantization — to meet latency expectations. We achieved this in the following phases: - Pre-Training. We trained Muse Glimmer on Muse Spark's outputs using logit distillation, leveraging a similar data mix as the teacher. - Mid-Training. We trained the model on longer-context, more agent-heavy data with richer reasoning traces, alongside organic data. - Post-Training. We combined supervised fine-tuning with a mix of on-policy distillation and reinforcement learning across general, reasoning, coding, and agentic domains. Muse Glimmer was evaluated under the standards set out in Meta's Advanced AI Scaling Framework and assessed for open-weight release across all relevant categories. Built for Agents: What Muse Glimmer Can Do Building effective agents requires key capabilities working together to achieve the user’s goals. Muse Glimmer is trained and evaluated across each of the following: - End-to-end Agentic Task Completion. Muse Glimmer achieves strong success rates on full-task benchmarks including DeepSearch QA, MCP-Atlas, 𝛕-Bench and SWE-Bench, which measure its ability to work within scaffolds, write and debug code, and resolve multi-turn requests from start to finish. - Reliable Tool Use. The model handles a wide range of function calls, invoking tools with precise schemas throughout extended workflows. - Multi-Step Reasoning. Muse Glimmer chains reasoning over long horizons, sustaining coherent plans across complex, extended workflows. - Failure Recovery. When a tool call fails or returns an unexpected result, the model is trained to diagnose the error and retry rather than halt. - Multimodal Input and Reasoning. Through a dedicated perception encoder, the model accepts interleaved text and images. This enables agents to interpret screenshots, charts, and documents alongside conversation. - Scaffold Compatibility. Muse Glimmer works across OpenClaw and other agentic orchestration patterns. - Controllable Effort. Muse Glimmer supports different reasoning strengths to select the right balance between quality and speed. - Multilingual. Muse Glimmer is trained on data from more than 100 languages. Performance We evaluated Muse Glimmer across a broad range of benchmarks to assess the diverse capabilities required for effective autonomous agent behavior. Compared with Gemma4-31B and Qwen3.6-27B, Muse Glimmer performs strongly for its size class on several widely used LLM benchmarks. For more detail about our evaluations, see our report. Optimized for Local Deployments A local agent is truly useful if it's fast enough to feel responsive. An agent that takes minutes to reply or plan its next step breaks the flow of real work. We applied two optimizations to make Muse Glimmer run at practical speeds on consumer hardware without sacrificing quality. Fitting the Model on Your Device. At full precision, a 30-billion parameter model would require over 55 GB of memory — far more than any consumer GPU offers. We use quantization techniques to compress the model's weights to approximately 4-bit precision, shrinking the language model to under 20 GB. This leaves enough headroom for the model's working memory (its "KV cache"), the perception encoder for image understanding, and the speculative decoding drafter to run simultaneously within a 24 GB or 32 GB envelope. We validated that this compression introduces minimal to no degradation on agentic tasks. Faster Generation Through Speculative Decoding. Language models normally generate text one token at a time, which can feel slow during long reasoning chains or multi-step tool calls. Muse Glimmer ships with a lightweight "drafter" model based on DFlash — a small companion network that proposes entire blocks of tokens at once. The main model then verifies these proposals in parallel, accepting correct tokens and correcting wrong ones. This technique lets Muse Glimmer generate text significantly faster than standard token-by-token generation while producing identical output quality. We provide quantized drafter versions to incur a smaller memory overhead in the release. The Result: We measure the speed of our K-Quant-17GB model alongside the quantized DFlash drafter on MacBook M4-Max, M5-Max and on a RTX-5090. The model is fast enough for fluid conversation and real-time agent interaction, all running entirely on your device. Get Started With Muse Glimmer Today Muse Glimmer is available now, and you can download the weights on Hugging Face. In the coming days, run it locally through partners like Ollama, LM Studio, and Unsloth, deploy it with edge frameworks including llama.cpp, ExecuTorch, and MLX, serve it at scale with vLLM and SGLang, or get started quickly through partners like Together AI, Fireworks AI, and OpenRouter. You can even customize it for your use case by leveraging PyTorch’s TorchTitan training feature to tune the model further. We're also working with our partners including AMD, Arm, Dell, Intel, and NVIDIA to optimize performance across devices. In addition, we’re releasing documentation so developers have the resources they need to get started and build responsibly with Muse Glimmer. This includes guidance on setting up custom scaffolds, so it's even easier to start building and deploying personal agents on day one. You can learn more and find resources to build on Meta's AI Developer Center. This work builds on Meta's long track record of open AI research, extending it into agentic AI and giving developers access to local agentic capabilities. As always, we welcome feedback from the community and can’t wait to see what developers build with this open weights model. Download the Model on Hugging Face Developer Documentation

3

Introducing Muse Glimmer

Simon Willison · original → · 8/10 · AI/work: Meta Muse Glimmer open agentic model
10th August 2026 - Link Blog Introducing Muse Glimmer (via) Meta are back in the open weights game! Muse Glimmer is a brand new 30B model under a clean Apache 2.0 license (a step up from the janky…

10th August 2026 - Link Blog Introducing Muse Glimmer (via) Meta are back in the open weights game! Muse Glimmer is a brand new 30B model under a clean Apache 2.0 license (a step up from the janky Llama licenses of old). They claim to have optimized it for exactly the kind of things I'm looking for in a local model: - End-to-end Agentic Task Completion. Muse Glimmer achieves strong success rates on full-task benchmarks including DeepSearch QA, MCP-Atlas, 𝛕-Bench and SWE-Bench, which measure its ability to work within scaffolds, write and debug code, and resolve multi-turn requests from start to finish. - Reliable Tool Use. The model handles a wide range of function calls, invoking tools with precise schemas throughout extended workflows. - Multi-Step Reasoning. Muse Glimmer chains reasoning over long horizons, sustaining coherent plans across complex, extended workflows. [...] Here's a pelican which I generated using LM Studio's 18.16 GB version of the model: I also tried it out with my llm-coding-agent plugin, running against a fresh checkout of Datasette with the prompt: how does auth work? Here's the response, at the end of a long transcript showing all of the tool calls it made to explore the codebase. I ran this using llm-lmstudio with this patch applied to upgrade it for compatibility with LLM 0.32. I really like this size of model, because if a machine has 32 GB of RAM or more (mine has 128GB) it leaves plenty of space for running other applications at the same time. Glimmer is a vision model, so I asked it to describe this image: llm -m lmstudio/meta/muse-glimmer -a https://static.inaturalist.org/photos/714731804/large.jpg 'describe image' Here's what I got back: The photograph shows a rocky, breakwater-style shoreline on an overcast day with a smooth, gray body of water and a faint dock/pier line in the soft-focused background. In the foreground two large brown pelicans, Pelecanus occidentalis, are perched on the jumbled gray-white stones. They have the species’ characteristic long, down-curved yellow-orange bills with a large throat pouch, long slender necks and mottled brown-gray plumage on the back and wings. The pelican on the left is turned slightly toward the camera and appears to be preening or resting its bill against its chest; a pale, whitish patch is visible on the crown and nape and a small crest of feathers is raised. The pelican on the right faces mostly forward/right, its head up, bill pointing down and to the right, with the same pale head markings and the barred, darker wing feathers clearly visible. Scattered among the rocks around the two pelicans are several much smaller dark birds — gulls/tern-like birds in muted brown-gray plumage. One dark bird sits on the far left on a rock, another brownish bird stands to the right of the right-hand pelican, a grayish bird with a reddish bill is in the lower right foreground, and a further small dark bird is at the extreme right edge of the frame. The overall light is flat and diffused, giving the water and sky a muted, almost monochromatic palette that contrasts with the textured rock and the detailed feathering of the pelicans. The composition places the two big birds as the dominant subjects, framed against the calm water and the low, rocky perch.

4

Golf championships under way at Enniscorthy

Wexford Local · original → · 7/10 · Local Wexford: golf championship at Enniscorthy
[image →]Avril Murphy receives her Captain’s Prize medal from Enniscorthy Lady Captain Helen D’Eathe. (Pic; Enniscorthy Golf Club) By Caomh Breen Allen Avril Murphy said Enniscorthy Golf Club are…
[image →]
Avril Murphy receives her Captain’s Prize medal from Enniscorthy Lady Captain Helen D’Eathe. (Pic; Enniscorthy Golf Club)

By Caomh Breen Allen

Avril Murphy said Enniscorthy Golf Club are ready to showcase the improvements made to their course when the Whirlpool Irish Senior Men’s & Women’s Amateur Close Championship gets underway on Tuesday morning.

Murphy has a dual role at this week’s championship, competing for the Women’s title while also helping to oversee the running of the event as Chair of the Management Committee at the Wexford venue.

“We’ve done major work with the course. We brought in a consultant two years ago to give us an overview of what we needed to improve on. The community of members started to come together and got behind all of this.”

“They’ve done a fantastic job around the place. The greenkeepers have been putting in fantastic work, they’ve brought our greens back to life, they’re absolutely fantastic.”

A total of 141 players will tee it up across the Women’s and Men’s championships, with the opening two rounds of Stroke Play qualifying taking place on Tuesday and Wednesday before focus switches to match play, culminating in the final rounds on Friday.

Having committed to hosting the championships more than two years ago, the event has provided a clear target for everyone at Enniscorthy to work towards.

Murphy will also get the opportunity to test those improvements herself as she bids to challenge for the Women’s title. Having won the Captain’s Prize at Enniscorthy earlier this year, she knows what is required to navigate the tree lined fairways.

“You need to keep your ball straight,” said Murphy.

“We’re narrow and we have a lot of trees. They say if you can play Enniscorthy well, you can play anywhere well. Your drive has to be accurate.”

“I’m looking forward to playing with the elite of golf from around the country.”

Murphy begins her championship at 12.40pm on Tuesday. Enniscorthy will also be represented by Connie Doyle and James Gilbert, with Murphy backing both local players to go well this week.

“Both Connie and James will fly the flag very well,” said Murphy.

“I have no doubt that they might even bring in a result.”

5

Show HN: Mcptoon – Token-efficient MCP CLI client

Hacker News · original → · 7/10 · AI/work: token-efficient MCP CLI for AI agents
MCP tool discovery costs 10,000+ tokens. mcptoon costs 350. One MCP client for every AI agent. Cross-platform. Zero dependencies. If this saves you tokens, please star the repo — it helps others…

MCP tool discovery costs 10,000+ tokens. mcptoon costs 350. One MCP client for every AI agent. Cross-platform. Zero dependencies. If this saves you tokens, please star the repo — it helps others discover it. English | 中文文档 | Report Bug | Request Feature Every MCP-enabled conversation burns tokens on syntax, not data: - Your agent connects to 5 MCP servers. Listing their tools: ~10,000 tokens of JSON. - Your agent calls 20 tools. Each returns 500-3,000 tokens wrapped in {"content":[{"type":"text","text":"..."}]} . - Total MCP overhead: 40,000-70,000 tokens before any actual thinking happens. On a 128K context window, that's 30-55% gone. Not on work. On syntax. mcptoon is a CLI client that connects to any MCP server (stdio or HTTP) and outputs TOON (Token-Optimized Object Notation) instead of JSON. | Operation | JSON tokens | mcptoon tokens | Savings | |---|---|---|---| | Tool discovery (96 tools) | ~2,000 | ~60 | 97% | | Tool result (structured data) | ~800 | ~350 | 56% | | Tool result (raw HTML/text) | ~1,000 | ~900 | 10% | Zero dependencies. Pure Python. 50KB. Works with every AI agent — Claude Code, Codex, OpenCode, Cursor, CatPaw, anything that runs shell commands. JSON (287 tokens) — what every other MCP client returns: [ {"name": "search_web", "description": "Search the web for information", "inputSchema": {"type": "object", "properties": {"query": {"type": "string", "description": "Search query"}, "num_results": {"type": "number", "default": 5}}, "required": ["query"]}}, {"name": "fetch_url", "description": "Fetch content from a URL", "inputSchema": {"type": "object", "properties": {"url": {"type": "string"}}, "required": ["url"]}} ] TOON (5 tokens) — what mcptoon returns: search_web fetch_url 98% reduction for tool discovery, 60% for full schema, zero information lost. pip install mcptoon Zero dependencies. 50KB. Python 3.10+. Windows, macOS, Linux. mcptoon init # Sample config: ~/.mcptoon/config.json mcptoon add fetch --stdio npx -y @modelcontextprotocol/server-fetch mcptoon manifest --toon # -> fetch:fetch mcptoon call fetch fetch '{"url":"https://example.com"}' --toon mcptoon call fetch fetch '{"url":"https://example.com"}' --json # when you need JSON | JSON | TOON | Why | |---|---|---| {"name":"search","count":3} | name:search|count:3 | Pipes replace braces + quotes + colon | [1, 2, 3] | 1 2 3 | Spaces replace brackets + commas | true / false | T / F | 1 char vs 4-5 | null | ∅ | 1 symbol vs 4 chars | "line1\nline2" | line1↲line2 | ↲ replaces escape sequence | {"a":{"b":[1,2]}} | a:b:1_2 | Recursive compaction | | Flag | What you get | Token footprint | |---|---|---| --toon | Compact notation, full semantics | 40-60% less than JSON | --compact | Tool names only, space-separated | 97% less than JSON | --json | Standard JSON (for scripts, CI) | Baseline | --raw | Raw response, no parsing | Full size | --head N | First N items only | Variable | --max-chars N | Hard truncate at N chars | Variable | --full | Disable the default 4000-char truncation | Full size | Set MCPTOON_AGENT_TYPE=claude and every call auto-selects --toon . | mcptoon | mcp-cli | mcporter | raw MCP SDK | | |---|---|---|---|---| | Token savings | 97% manifest, 40-60% results | 0% | 0% | 0% | | Works with all agents | yes (Claude Code, Codex, OpenCode, Cursor, any) | Claude only | Claude only | varies | | One config for all agents | yes | no | no | no | | Output formats | TOON + JSON + compact | JSON | JSON | JSON | | Dependencies | 0 | 5-20 | npm | 3-8 | | Dangerous-op blocking | yes | no | no | no | | Usage tracking | yes (local) | no | no | no | | Schema cache | yes (5min) | no | no | no | | Install size | ~50KB | ~50MB+ | ~30MB | ~10MB | | Platform support | Windows, macOS, Linux | Linux/macOS | macOS | varies | mcptoon is a CLI tool. If your agent can run shell commands, it can use mcptoon. | Agent | How to use | |---|---| | Claude Code | Write mcptoon commands in SKILL.md files | | Codex (OpenAI) | Add mcptoon to AGENTS.md | | OpenCode | Use mcptoon in custom commands | | Cursor | Add mcptoon to .cursorrules | | CatPaw | Write mcptoon commands in skill files | | Any agent | If it runs shell commands, it can call mcptoon | Configure MCP servers once in ~/.mcptoon/config.json . Every agent shares the same servers, the same tools, the same token savings. export MCPTOON_AGENT_TYPE=claude # auto-select --toon # In ~/.claude/skills/mcp-tools/SKILL.md Search the web: mcptoon call exa search '{"query":"AI news"}' List available tools: mcptoon manifest --toon Fetch a URL: mcptoon call fetch fetch '{"url":"https://example.com"}' # In AGENTS.md or system prompt Use mcptoon to call MCP tools. It saves 60% tokens vs JSON. - List tools: mcptoon manifest --toon - Call a tool: mcptoon call <server> <tool> '{"args":"here"}' --toon from mcptoon.client import MCPClient from mcptoon.output import toon with MCPClient(stdio=["npx", "-y", "@modelcontextprotocol/server-fetch"]) as c: tools = c.list_tools() print(toon(tools)) # compact TOON result = c.call_tool("fetch", {"url": "https://example.com"}) print(toon(result)) from mcptoon.router import register @register("my-database", "db") def handle_db(tool, args): if tool == "query": return {"rows": my_db.execute(args["sql"])} return None # falls through to MCP # stdio (any npx MCP server) mcptoon add fetch --stdio npx -y @modelcontextprotocol/server-fetch mcptoon add github --stdio npx -y @modelcontextprotocol/server-github # HTTP mcptoon add myapi --http http://localhost:3001/mcp --header "Authorization: Bearer xxx" Config lives at ~/.mcptoon/config.json . Project-level override at ./.mcptoon.json . mcptoon blocks operations that match dangerous patterns (delete , drop , purge , wipe , kill , etc.) unless you pass --destructive . $ mcptoon call db delete_table '{"name":"users"}' Error [CONFIRMATION_REQUIRED]: Dangerous operation needs confirmation $ mcptoon call db delete_table '{"name":"users"}' --destructive # runs $ mcptoon usage Total calls: 142 Success rate: 138/142 Tokens (est): 84,200 By server: fetch 89 github 53 Stored locally at ~/.cache/mcptoon/usage.json . Never transmitted. src/mcptoon/ ├── cli.py # CLI entry + arg parsing ├── client.py # MCPClient — stdio + HTTP transport ├── router.py # Tool routing, custom handlers, safety checks ├── config.py # Server config ├── manifest.py # Tool discovery with cache ├── output.py # TOON / JSON / compact rendering ├── cache.py # Schema cache (5-min TTL) ├── usage.py # Local usage tracking └── errors.py # Structured error envelopes ~1,700 lines total. Zero third-party imports. - No telemetry. No analytics, no crash reports, no phone-home. - No credential storage. API keys pass through from your config or env vars. - No dependencies. Pure Python stdlib. No supply chain to audit. Found a vulnerability? Email security@activeing123.github.io . See SECURITY.md. Apache 2.0. See LICENSE and NOTICE. git clone https://github.com/activeing123/mcptoon.git cd mcptoon pip install -e . --no-build-isolation pip install pytest pytest-cov python -m pytest tests/ -v # 98 tests, 0.09s Zero dependencies is a hard rule. New features need tests. See CONTRIBUTING.md. mcptoon is an independent third-party MCP client. Not affiliated with Anthropic. Found this useful? Star the repo to help others find it.

6

Show HN: Needle2: 14MB agentic LLM for phones, wearables, smart home and robots

Hacker News · original → · 7/10 · AI: small efficient agentic LLM for mobile devices
Today we release Needle 2: an open 45M-parameter model for tool calling, device use and structured extraction. The whole model is a single 14MB binary that runs a full session in 28MB of RAM. It is…

Today we release Needle 2: an open 45M-parameter model for tool calling, device use and structured extraction. The whole model is a single 14MB binary that runs a full session in 28MB of RAM. It is built on our Simple Attention Network findings, compressed to CQ2-bit with Cactus Quants, and baked into its own engine. On the tool call and mobile device use benchmarks, Needle 2 trades wins with other small models like FunctionGemma 270M, LFM2.5 230M and Apple FM, at 5× to 70× smaller, and 2 bits against their f16. Needle hits 500 tokens/sec decode speed on a Raspberry Pi 5, between 400–1,500 tokens/sec on VR devices like Meta Quest 3S and Apple Vision Pro, and ranges 300–700 on sub-$200 phones such as the Samsung A-Series. With a peak session RAM around 28MB, Needle runs on newer microcontrollers like ESP32-S3. The Playground lets you test Needle for wearables, robots, smart homes, phones, and automotive. Needle is licensed under Apache 2.0, with weights on Hugging Face; the repo gets you running. - 45M - Params - 800+ tok/s - Pi5 prefill - 500+ tok/s - Pi5 decode - CQ2-bit - Compression - 14 MB - File size - 28 MB - Session RAM Our Bet Bringing On-Device AI to <$200 Devices: Edge AI has lately meant Macs and PCs, but the edge is mostly cheap hardware: over 21 billion connected IoT devices against roughly 1.5 billion PCs, and in emerging markets most phones ship under $200. Count budget phones, Raspberry Pis, microcontrollers, wearables, small robots like Reachy Mini, and connected home devices, and roughly four in five edge devices cost under $200. That is the hardware Needle targets: no GPU, no NPU, a few hundred MB of RAM. Function Call & Device Use: Turning on a light does not need a frontier model. A watch, a home, a robot: each already exposes its abilities as functions with typed parameters, so the only hard part is mapping a messy sentence onto them: which function, with which values. Framed that way, the problem needs no world knowledge and no open-ended prose, which is why 45M parameters suffice where chat needs billions. That smaller formulation is the bet everything else follows from. Extraction & Structured Outputs: The schema is the interface, and the same formulation covers documents: a schema plus a paragraph returns typed fields, an enum field is a classifier, an array field collects a list in one call. We enforce this with a contract, not a convention: every turn is answered with a call envelope, the empty call is the refusal, and a byte-level grammar compiled from the declared schemas constrains every token. The grammar carries the syntax, so all 45M parameters go to choosing functions and grounding arguments in the user's words. Edge-Cloud Collaboration: No small model covers everything, so Needle says so instead of guessing: every response carries a learned confidence score, and off-topic requests return the empty call. Above your threshold, act; below it, re-ask or escalate to the cloud. Most device requests are routine control, so escalation stays rare and the default path stays private, instant, and free. Lossless 2bit Quantization: Small models break under post-hoc quantization, so we never quantize post-hoc: Needle 2 trains against Cactus Quants from pretrain through post-train, weights, activations, and KV cache alike. The 2bit model you deploy is the model that was trained. That is what fits 45M parameters into 14MB with nothing lost on our battery. Co-designed Model & Inference: Every architectural choice was benchmarked on the target hardware before it earned its parameters, and the deliverable is the pair, not the weights: a single dependency-free C++ binary that probes the CPU at startup and picks its kernels, with the model, tokenizer, and grammar compiler sealed inside. One artifact runs from Cortex-M to x86 to WebAssembly. There is nothing to install and nothing to download. Fine-tune on your Mac/PC: Every product has its own tool vocabulary, and a 45M model is small enough to retrain where it runs: the repo and python package tune and test on your own computer in minutes to a few hours. Ship a Needle that speaks your device's tools, not a generic assistant. Production Needle is production-ready for products that require a minimal RAM footprint, low latency, privacy, and offline reliability. Pebble - the pioneer of the modern wearable industry - runs it locally in the Index 01 app to turn spoken requests into actions without depending on a network connection. The Pebble Index Ring has no screen. So when you speak to it, the action just has to happen, every time, with or without internet connection. We run Cactus Needle locally in the app, instead of relying on the cloud. The model's footprint is tiny and the performance never lets us down. Architecture Needle 2 is pretrained on a proprietary 115B-token corpus and post-trained on 38B tokens with compact reasoning traces and careful dataset distribution design. For scale: LFM2.5-230M was pretrained on 19 trillion tokens, roughly 120× Needle's total, and the evaluation below shows the two trading wins. Each component exists to buy capability without buying bandwidth. The Hadamard MLP replaces the usual dense up-and-down projections with a fixed Walsh transform and learned diagonals, so the channel mixing that dominates a small model's weight reads costs almost no parameters at all. The engram moves world knowledge out of the stack into hashed n-gram tables that are read a few rows per token: capacity that is nearly free at decode time, which matters on devices where every megabyte read from flash is latency and battery. The multi-lane residual streams give a 27-layer, 512-wide network the routing flexibility of a much wider one, at the cost of a few dot products per layer rather than more attention or MLP volume. The memory system is designed backwards from fixed-RAM devices. Attention uses a 256-token sliding window so the KV cache is bounded no matter how long a session runs, and the system prompt and tool declarations are pinned as permanent sinks so the one thing a tool-calling model must never forget—its tools—is structurally unable to be evicted. The cache itself is trained with QAT, and weights are stored in Cactus Quants at a mixed bits per weight averaging 2bit. The result is that quality decisions and deployment decisions stay decoupled: one trained model, specialized to whatever precision and window a target device can afford. The engine earns its speed from what it refuses to compute. Weights never decompress into RAM: the 2-bit codes are expanded inside vector registers, fused into integer dot products, so resident memory stays at blob size and the arithmetic path is int8 end to end—activations, KV cache, and the lane routing tables alike. The grammar is an optimization, not just a guarantee: because the matcher knows which tokens are legal before the logits exist, the engine computes output scores only for candidate rows, skipping up to 98% of the vocabulary projection on structural tokens, and skips it entirely on steps whose output is already forced. One universal binary probes the CPU at startup and self-selects its kernel tier—SDOT, NEON, AVX2, RISC-V vectors, wasm SIMD, or scalar—and the thread pool spins through the short serial sections of a token instead of sleeping, which alone nearly doubled decode. None of this changes a single output: every trick is either exact or validated token-for-token against the reference path. All of it is ultimately an energy argument. On device silicon, moving a byte out of flash or DRAM costs orders of magnitude more than a multiply-accumulate, so the budget that matters is FLOPs per token and bytes per token together. The architecture cuts the first: a conventional transformer of Needle's width and depth spends 164 MFLOPs per token, and even one squeezed down to Needle's parameter count spends 87, because every parameter it owns must be exercised through a matmul. Needle spends 70, and keeps a fifth of its parameters as gathered memory that costs no arithmetic at all. The binary cuts the second, as the engine section showed: nothing rematerializes, the arithmetic stays int8 end to end, and the grammar prunes compute outright, so decoding a token reads at most the 14MB blob once, and on structural tokens meaningfully less. This is what battery life is made of. Even on a high-end phone, an always-on assistant lives inside a power budget; every MFLOP is milliwatt-hours, and Needle spends 7× to 85× fewer of them per token than the models it is benchmarked against. Compute per token | Model | Params | Matmul-active | MFLOPs / token | |---|---|---|---| | Needle 2 | 45M | 35M | 70 | | Same-shape transformer, dense MLP | 82M | 82M | 164 | | Transformer at matched params | 43M | 43M | 87 | | LFM2.5 230M | 230M | 230M | 460 | | FunctionGemma 270M | 270M | 270M | 540 | | Apple FM | ~3B | ~3B | ~6,000 | Bounded session memory is what puts microcontrollers in reach. Because the sliding window caps state, Needle 2's RAM is a deterministic 28MB ceiling, not a curve that grows with conversation length. That fits MCU-class parts with external RAM, such as ESP32-P4 with 32MB of PSRAM, or STM32H7 and NXP i.MX RT boards with SDRAM. The engine compiles single-threaded for bare metal and ships as a static library for Cortex-M4, M7, and M55. Evaluation We evaluate on five public function-calling benchmarks: Google's Mobile Actions, DroidCall, the Seal-Tools in-domain and out-of-domain tests, and BFCL v4 single-turn. Scoring is ordered strict exact match: a row passes only if the function names, the call order, and every argument value match. All Needle 2 numbers are measured end-to-end through the shipped C++ engine in its production configuration: CQ2-bit weights, tool retrieval on, and the 256-token sliding KV window. Nothing is relaxed for benchmarking; the numbers reflect the exact engine a device runs, window eviction included. Baselines run the released checkpoints under vLLM at full context, and Apple FM runs on-device. Two asymmetries make this comparison hard, and we state both upfront. Precision: the baselines stay at f16 deliberately, because conventional post-training quantization to 2 bits collapses models that were never trained for aggressive compression, while Cactus Quants is baked into Needle's training from the ground up. That skew favors the baselines. Scope: Needle is trained specifically for agentic tool calling and nothing else, while every baseline is a general language model carrying chat, prose, and world knowledge alongside its tool calling. That skew favors Needle. There is no clean way to level both at once, so we do not try. The tables answer one narrow question: which model executes tool calls correctly within an on-device budget. We accept the skew; it still paints the picture we intend. Mobile Actions (961 rows) | Model | Accuracy | Name acc. | Non-empty | 1-call | 2-call | |---|---|---|---|---|---| | LFM2.5 230M (f16, vLLM) | 69.1 | 93.0 | 98.9 | 76.1 | 55.0 | | FunctionGemma 270M (f16, vLLM) | 64.0 | 87.3 | 98.9 | 73.0 | 46.2 | | Needle 2 (CQ2-bit) | 63.7 | 98.3 | 99.4 | 71.3 | 48.4 | | Apple FM (on-device) | 57.6 | 94.2 | 95.5 | 64.5 | 43.8 | DroidCall test split (200 rows) | Model | Accuracy | Name acc. | Non-empty | 1-call | 2-call | |---|---|---|---|---|---| | FunctionGemma 270M (f16, vLLM) | 17.5 | 37.5 | 59.5 | 22.7 | 0.0 | | Needle 2 (CQ2-bit) | 17.0 | 36.5 | 47.5 | 22.1 | 0.0 | | LFM2.5 230M (f16, vLLM) | 11.0 | 21.5 | 22.5 | 14.3 | 0.0 | Seal-Tools in-domain (700 rows) | Model | Accuracy | Name acc. | 1-call | 2–3-call | 4+-call | |---|---|---|---|---|---| | Needle 2 (CQ2-bit) | 32.6 | 64.9 | 63.0 | 21.8 | 14.6 | | LFM2.5 230M (f16, vLLM) | 26.9 | 45.4 | 54.5 | 17.1 | 10.4 | | FunctionGemma 270M (f16, vLLM) | 16.3 | 56.0 | 47.0 | 4.5 | 2.1 | Seal-Tools out-of-domain (654 rows) | Model | Accuracy | Name acc. | 1-call | 2–3-call | 4+-call | |---|---|---|---|---|---| | Needle 2 (CQ2-bit) | 28.7 | 58.7 | 56.4 | 27.1 | 15.4 | | LFM2.5 230M (f16, vLLM) | 17.0 | 35.0 | 42.6 | 13.7 | 9.8 | | FunctionGemma 270M (f16, vLLM) | 15.6 | 48.9 | 50.0 | 11.0 | 6.3 | Needle was not trained for general function calling: its corpus is consumer device actions—smart home, mobile, wearables, TV, car—plus structured extraction, and BFCL's general-purpose and enterprise API surfaces, including the Java and JavaScript SDK categories, sit entirely outside that distribution. It extrapolates nonetheless: on Python simple calls it lands within a point of FunctionGemma, a model six times larger trained for exactly this task, and it keeps a 93.4 well-formed rate across all 3,641 rows. The gap concentrates where its training data has never been: Java, JavaScript, and the parallel multi-call categories. BFCL v4 single-turn (3,641 rows) | Category | Apple FMon-device | LFM2.5 230Mf16 · vLLM | FunctionGemma 270Mf16 · vLLM | Needle 2CQ2-bit | |---|---|---|---|---| | Simple | 73.3 | 63.2 | 48.1 | 40.8 | | — Python | 86.8 | 85.5 | 62.3 | 61.2 | | — Java | 67.0 | 48.0 | 38.0 | 29.0 | | — JavaScript | 66.0 | 56.0 | 44.0 | 32.0 | | Multiple | 84.0 | 78.5 | 60.0 | 57.0 | | Parallel | 65.0 | 64.0 | 36.5 | 30.0 | | Parallel multiple | 52.0 | 51.5 | 30.5 | 22.5 | | Live simple | 70.5 | 45.0 | 33.7 | 36.8 | | Live multiple | 45.9 | 47.8 | 25.2 | 27.9 | | Live parallel | 50.0 | 43.8 | 18.8 | 25.0 | | Live parallel multiple | 58.3 | 45.8 | 25.0 | 29.2 | | Relevance | 100.0 | 68.8 | 81.2 | 81.2 | | Irrelevance | 28.3 | 77.7 | 72.1 | 60.8 | | Overall | 61.7 | 60.8 | 46.1 | 42.6 | | Well-formed rate | 95.0 | 94.2 | 100.0 | 93.4 | Explore Needle 2

7

🎙️ How I AI: Build an AI code review bot in 30 minutes + Claude Code for normal people

Lenny's Newsletter · original → · 7/10 · AI/work: building AI code review bot with Claude
[image →]Build an AI code review bot in 30 minutes with Vercel EveListen now on YouTube • Spotify • Apple Podcasts[image →]Brought to you by:WorkOS—Make your app Enterprise Ready, with SSO, SCIM,…

Build an AI code review bot in 30 minutes with Vercel Eve

Listen now on YouTube • Spotify • Apple Podcasts

Brought to you by:

WorkOS—Make your app Enterprise Ready, with SSO, SCIM, RBAC, and more

Claire built an AI agent that reviews pull requests, scores their risk, auto-approves the safest ones, and sends anything questionable to Slack. In this episode, she breaks down how she built the entire thing in one Codex session using Vercel Eve, and why letting AI review AI-generated code may be a lot less risky than it sounds.

Biggest takeaways:

  1. You don’t need a human to review every AI-generated PR. That may sound reckless, but it’s increasingly looking like the smarter operating model. Intercom has already shown this can work at scale: PRs approved by its AI system move five times as fast as human-reviewed ones and have a lower revert rate. In other words, the AI-reviewed code isn’t just shipping faster—it’s less likely to need fixing after it reaches production.

  2. The key is having a clear way to separate changes that can skip human review from ones that can’t. The risk model used here scores each PR across six dimensions: the size of the change, its blast radius, how easily it can be reversed, its data and security implications, its operational impact, and whether tests and CI have actually been completed. Anything below 24 points is classified as low risk and cleared by the agent; anything above 64 goes straight to a human for review. The exact numbers matter less than turning a vague judgment call into a repeatable system.

  3. Vercel’s Eve may be the fastest way to deploy a serious internal AI agent across Slack and GitHub without spending weeks building infrastructure. It handles the annoying plumbing—connectors, refresh tokens, sandboxing, and routing across channels—so the actual work becomes writing instructions and skills in Markdown instead of babysitting OAuth flows.

  4. A useful internal agent can now be built and deployed in a single Codex session, starting with a prompt that’s only a couple of sentences long. In this case, the initial ask was essentially: build a GitHub bot that waits for checks to pass; grades each PR as low, medium, or high risk; and automatically approves the low-risk ones. Everything after that was steering and refinement, not a giant up-front specification.

  5. Browser use removes much of the configuration tax that makes agent setup feel harder than it should. Creating a Slack bot and GitHub app manually normally means clicking through endless permission screens, choosing scopes, and managing tokens. Codex handled almost all of that through the browser. The human’s job was mostly to click “save” and complete 2FA. What usually takes hours took minutes.

  6. SOC 2 compliance and automatically approved PRs are not inherently at odds. The important part is making the process legible: the risk model needs to be reflected in the company’s code-review and security policies, every decision needs to be logged, and the resulting audit trail needs to be easy to query and defend. The security team’s role is to help design the right framework, not simply block automation because it feels unfamiliar.

  7. The operational design matters just as much as the underlying technology. Merge Mommy doesn’t actually merge anything. Instead, it posts a gray check in GitHub as a signal, then sends a Slack message with the risk score and a note saying the PR is ready to approve and merge. That small handoff preserves human accountability for the final action while eliminating most of the cognitive work involved in reviewing a routine change.

  8. Evals are what keep internal agents trustworthy after the novelty wears off. Intercom logs every PR review its agent produces and then has an engineer assess whether the score and recommendation were correct. That’s the same discipline strong teams already bring to customer-facing AI products. Internal agents may feel less visible, but when they touch something as important as the codebase, they need the same protection against regressions.

  9. The surprising thing about building an Eve agent is how little “building” is actually involved. The full instructions for Merge Mommy fit on roughly a page: a few paragraphs, a handful of bullets, and a short skill file. There isn’t much framework-specific magic to learn. The core skill is simply being able to explain, clearly and precisely, what the agent should do.

Blog from this episode:

How I Built ‘Merge Mommy’: My AI Bot for Auto-Reviewing Pull Requests with Vercel Eve: https://www.chatprd.ai/how-i-ai/merge-mommy-vercel-eve-ai-bot-for-auto-reviewing-pull-requests

Claude Code for normal people: skills, voice mode, and how to collaborate with AI

Listen now on YouTube • Spotify • Apple Podcasts

Brought to you by

  • Bolt.new—Turn your idea into a real product

  • Hyperagent—Deploy fleets of agents that handle real work

Grace Clarke is an AI educator and former marketing consultant who rebuilt her entire service business around Claude. In this episode, she walks through the three Claude skills that run her business: an hourly client pipeline, a proposal builder, and a voice guide that teaches Claude how she thinks. She also shows how she replaced Gmail with her own Claude-powered inbox, why she believes intent engineering matters more than prompt engineering, and how non-technical people can start building useful AI workflows without writing code.

Biggest takeaways:

  1. You don’t need to be technical to build a business workflow with Claude. You need to be clear about what’s broken. Grace started by opening Claude Code and talking for a few minutes about everything driving her crazy: too many emails, client communication that didn’t feel warm enough, and 20 hours of admin every week. Claude helped turn that into an operating system for her business.

  2. Intent engineering matters more than prompt engineering. Grace doesn’t spend time crafting the perfect prompt. She explains the problem, describes the outcome she wants, and asks Claude to come back with a proposal. Her philosophy is simple: the burden of figuring out what to do next shouldn’t always fall on the person. Give Claude enough context, and it should be able to study how you work and bring you strong ideas to react to.

  3. Grace’s voice guide isn’t really a writing guide. It’s closer to a “think like Grace” file. It captures how she makes decisions, what she believes about teaching versus consulting, what good communication sounds like, and even the kinds of LinkedIn posts she never wants to resemble. Whenever she sees something that makes her cringe, she sends Claude a voice note and updates the guide.

  4. Doing all of your email inside Gmail means almost none of that work compounds. Every response contains useful context: how a client communicates, what tone works with them, what they care about, and the history of the relationship. But if that context stays buried in Gmail, the AI never gets smarter from it. Grace moved her inbox workflow into Claude so every interaction can become useful context for the next one.

  5. Grace has developed a surprisingly elegant way to move between Claude Code and Cowork. She uses Claude Code when she wants something proactive that will go figure things out. Then, when she wants a more visual and approachable environment, she has Claude Code create a Markdown handoff file and drops it into a new Cowork session. The result is a simple way to carry context from one environment to another without starting over.

  6. The biggest barrier to adopting AI may be less about fear and more about muscle memory. Grace used to give students carefully engineered prompts so they could experience an immediate win. Her students told her that approach was actually making things harder. What helped more was building the habit of reaching for AI in the first place. Now she’ll do things like send a Slack reminder telling students to screenshot whatever they’re working on and drop it into Claude. No perfect prompt, no required outcome—just practice.

  7. Skill files may be one of the most underrated ideas in agentic AI. Grace runs much of her business around three of them: a pipeline operator, a proposal maker, and a voice guide. Each is essentially a documented set of instructions Claude can reuse. She teaches students to build them the same way they would train a new employee: explain the job, walk through how it should be done, correct mistakes, and keep improving the instructions over time.

  8. One of the best ways to sell someone on AI is to show them what it can do before explaining it. Grace gives new clients a password-protected, branded, interactive welcome experience she built in Claude. It’s part onboarding, part proposal, and part demo. By the time a client enters the password and starts clicking around, they already understand what makes this new way of working different.

Blog and detailed workflow walkthroughs from this episode:

Grace Clarke’s Claude Workflows for Business Automation and Rebuilding Gmail

https://www.chatprd.ai/how-i-ai/claude-workflows-for-business-automation-and-managing-gmail

↳ Automate Client Proposals and Onboarding with a Claude “Pipeline Operator”

https://www.chatprd.ai/how-i-ai/workflows/automate-client-proposals-and-onboarding-with-a-claude-pipeline-operator

↳ How to Rebuild Your Gmail Inbox Inside Claude to Manage Email

https://www.chatprd.ai/how-i-ai/workflows/how-to-rebuild-your-gmail-inbox-inside-claude-to-manage-email

↳ Create an Automated Workout Tracker with a Simple Claude Voice Note

https://www.chatprd.ai/how-i-ai/workflows/create-an-automated-workout-tracker-with-a-simple-claude-voice-note


If you’re enjoying these episodes, reply and let me know what you’d love to learn more about: AI workflows, hiring, growth, product strategy—anything.

Catch you next week,
Lenny

P.S. Want every new episode delivered the moment it drops? Hit “Follow” on your favorite podcast app.

8

Claude Code for normal people: skills, voice mode, and how to collaborate with AI

Lenny's Newsletter · original → · 7/10 · AI/work: Claude Code for building SaaS tools
Grace Clarke is an AI educator and former marketing consultant who taught herself Claude Code earlier this year and built a curriculum out of the process. She now runs her entire service business on…

Grace Clarke is an AI educator and former marketing consultant who taught herself Claude Code earlier this year and built a curriculum out of the process. She now runs her entire service business on tools she’s built with Claude, including a pipeline operator, a proposal maker, and a Gmail replacement she created in under 30 minutes, and teaches individuals and teams to do the same.

Listen or watch on YouTube, Spotify, or Apple Podcasts

What you’ll learn:

  1. How to build an hourly pipeline in Claude that moves clients through your process automatically

  2. Why Grace ditched traditional proposals for password-protected, interactive HTML documents built in Claude

  3. How she uses a “voice guide” skill file so every Claude output sounds like her, not like AI slop

  4. The two-step forcing function she teaches non-technical clients to build the muscle of opening Claude

  5. Why she started building in Claude Code, then handed the work off to Cowork via a Markdown session file

  6. How she replaced Gmail entirely with a custom inbox

  7. Why she teaches “intent engineering” instead of prompt engineering, and what that looks like in practice

  8. How she uses Claude on her phone, on walks, to track workouts and manage plants alongside client work


Brought to you by:

Bolt.new—Turn your idea into a real product

Hyperagent—Deploy fleets of agents that handle real work

In this episode, we cover:

(00:00) Grace’s background and why she started building with Claude

(04:48) The pipeline operator: what it is and how it runs her business every hour

(08:48) Building the muscle memory to use AI

(12:02) What goes into building a skill file (voice guide, proposal rules, versioning)

(13:50) How she built her proposal maker

(16:15) The voice guide: teaching Claude how she thinks, not just how she writes

(21:22) Live demo of the custom Gmail replacement built in Cowork

(30:44) Workout tracking, plant photos, and tiny daily Claude habits

(34:51) The biggest misconception holding people back from adopting AI

(38:36) What Grace does when Claude is not giving her what she wants

(40:38) Claude builds a proposal for Claire in real time

Tools referenced:

• Claude: https://claude.ai

• Claude Code: https://claude.ai/code

• Netlify: https://www.netlify.com

• Google Forms: https://forms.google.com

• Google Sheets: https://sheets.google.com

• Google Cloud (for service accounts and custom connectors): https://cloud.google.com

Other reference:

• Stratechery by Ben Thompson: https://stratechery.com

Where to find Grace Clarke:

LinkedIn: https://www.linkedin.com/in/gracegclarke/

X: https://x.com/graceclarke

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

Items scoring 7/10 or above from 11 sources, scored by claude-haiku-4-5-20251001 on relevance to my interests. At most 3 per source.

Scoring categories & sources
  1. Local Wexford or South East Ireland news
  2. Irish or EU-wide affairs affecting citizens broadly: elections, new laws or policy being debated, cost of living, education — especially impacts on mid-life adults or teenagers. Never courts/crime stories.
  3. Irish news on a topic relevant to my interests
  4. Work and tech topics: networking, AI, Kubernetes, platforms, SaaS
  5. AI news including critical or anti-AI perspectives
  6. Gaming: PC gaming, indie gaming, retro gaming
  7. General interests: gardening, woodwork, cycling, fitness, travel
  8. Comics

Sources: Breaking News Ireland, Wexford Local, Hacker News, r/gaming, r/pcgaming, r/antiAI, r/indiegaming, Lenny's Newsletter, One Useful Thing, Newcomer, Simon Willison

Comics

Size and Lifespan

XKCD · view →
With their 13 years recording and performing together and two humans worth of mass, the White Stripes are sandwiched neatly between gray wolves and blue whales.

With their 13 years recording and performing together and two humans worth of mass, the White Stripes are sandwiched neatly between gray wolves and blue whales.