daily

2026-08-23
1

Why your local LLM feels dumber than it is

Hacker News · original → · 8/10 · AI: practical LLM inference implementation issues
Quick Introduction We have all been on forums, chats, reddit, discord, youtube, or somewhere and heard “Oh! Model XYZ is AMAZEBALLZ!zomgwtfbbq” then downloaded it (or more likely, some quantized…

Quick Introduction We have all been on forums, chats, reddit, discord, youtube, or somewhere and heard “Oh! Model XYZ is AMAZEBALLZ!zomgwtfbbq” then downloaded it (or more likely, some quantized form of it) and said “eww… This sucks!” This post is going to be a rather technical series of experiments to demonstrate the impact of implementation-specific hazards with inference. I will be using the term “reference implementation” to describe the lab that published and offers first-party hosting of their models and posts original benchmark claims. Their hardware will be different than yours. Their software will be very different than yours. And the comparisons in this post are not going to be running some 2.58-bit-gguf-in-ollama with a couple test prompts. I am intentionally glossing over entire emerging fields of study, mountains of research papers and lit review to make this more approachable for you the reader. Don’t nit pick my oversimplifications or I will make you read the really long unpleasant version with math. Your local implementation sucks. But that’s ok, because everyone else’s does too. Every single instance of hardware and software running an LLM today is a little bit different. or a lot different when it comes to some cases. The average home lab user might be mixing multiple different generations of GPU. The chips on those have different instruction sets. Those instruction sets will implement and execute math to calculate your next token differently from any other person, even when running the same exact weights. So that begs the first question: How much does your particular setup suck? Turns out there are a number of different ways to go about measuring that. The practical approach is straight forward. Run standard benchmarks. A variety of them. terminal bench, hle, SWEthis, HELLAthat, MMLU-whatever… take your pick. Just make sure its representative of your actual workload/use case. Do not crank temperature to zero and paste in 3 test prompts then call it good/bad. Zero-shot tests are not a good analog of most agentic tasks. You need long-context tool-calling and domain specific knowledge evaluations to figure out where your setup is weak when running the same weights as somebody else replicating those same benchmarks. But the purely mathematical answer is where my focus is going to begin because as @wendell said: Math is Math! “Logits” are the models scores for each possible next token. They are normalized into probabilities, passed through the configured sampler, and converted back into text by the detokenizer to generate THE→NE→XT→TOK→EN during decode. A side note about sampler settings: the model card on HF usually specifies exactly what sampler settings (and chat template) you should be using. temp 1.0, top-p 0.95, etc. it varies by model so make sure you are using the right ones. btw, setting temp too low is why your qwen is sitting there looping unable to escape its THINK output. You’re welcome, glad I could fix that for you. When the next token probability changes enough, THE→NE→XT becomes THE→NE→W→DAY… And while those small changes might be fine, odds are that’s the beginning of the niggling sensation in the back of your mind that something feels off. Some of you may have heard the term KLD before, or KL Divergence. Don’t worry, I won’t make you do any math or flood your brain with tables of very small decimal numbers. But just in case you wanted the simple version: convert the output logits into a probability distribution, and measure how far that distribution has moved from a chosen baseline. Lower KLD means closer to that baseline, not automatically ‘smarter’. KLD is also directional, so the order of the two distributions matters. A word of caution: Don’t get suckered in by impossibly low KLD claims on a quant HF model card. It is impossible to interpret a number unless the author discloses the reference checkpoints and full runtime environment, evaluation text, calibration data, context lengths, sampled positions, KL direction, any vocabulary truncation, and how the measurements were aggregated. The methodology matters as much as the number and plenty of people get it wrong. What the hell is vllm doing? Now, we need to take a brief field trip down what the giant stack of software is doing on your inference engine to understand where some of those sources of divergence come from. At every step of this oversimplified diagram are components that can be configured or changed based on your specific hardware/software footprint, model, quant, tensor shape, etc. The nightly VLLM container image I snagged had 734 (252 uv/pip Python) packages in it. That’s 734 codebases each with their own bugs and undocumented idiosyncrasies. The path your specific implementation takes through that mountain of code will be distinct. Test 1: Precision Benchmarking Attention Backends Lets start with one piece of that inference flowchart. During prefill (prompt processing) there are a several attention backends your inference engine will select from. This impacts both speed and precision of prefill, while requiring different cuda kernels for every GPU family / SM compute capability 1.3. The CUDA platform — CUDA Programming Guide . Lets test them and compare. (I’m really very sorry, I had to…)I started with the official BF16 checkpoint of Qwen3.6-27B on an RTX PRO 6000 Blackwell GPU at tensor parallelism 1. The KV cache was BF16, with no weight/activation or KV-cache quantization. The software was a pinned nightly vllm build. I used eager execution, disabled CUDA graphs, prefix caching, and MTP, and used 2k-token chunked prefill. Qwen3.6-27B is dense, not an MoE, but it is still a hybrid model. 64 layers repeat in a pattern of three Gated DeltaNet/linear-attention layers followed by one full-attention layer. Only those 16 full-attention layers use the selectable attention backend in this experiment; the Gated DeltaNet path remained fixed. The workload replayed here is “Prompt 2”, a roughly 100k token context captured from a real Turnstone lab workstream containing multiple tool calls and real work products. It was selected to resemble what a local agent actually does rather than a synthetic needle-in-a-haystack test. And maybe more importantly, it doesn’t appear in any benchmark or training dataset in the wild today. Nobody could have benchmaxed for this, or calibrated their quant to accommodate it. There are three available full attention backends to select from in vllm for this workload: FlashAttention 2, Flash Inference, and Triton Attention. This was the only change made between executions, the rest of the hardware and software stack remained stable. I also performed a same-backend cross-GPU repeatability control. For this graph, I captured the full-vocabulary logits in BF16 every 32 prompt tokens. Distribution comparisons such as KLD were calculated afterward in FP64 from those stored logits. Top-1 agreement is whether the token with the highest logit, the greedy argmax, was the same. All three backends were evaluated against the same forced token history. A “top-1 flip” therefore means a backend would have chosen a different greedy next token at that position. We did not let that choice alter the remaining history. This keeps the mathematical comparison controlled, but it does not show how far an unconstrained generation would branch or whether a tool call would eventually fail… that comes in test 2 ;D The following graph shows % of sampled logits resulting in token flips: For the first several thousand tokens, every run of the model agreed about what the next token was going to be regardless of backend. Then in later portions of the prompt, backends began disagreeing. Triton was selected as the baseline to simplify upcoming quantization chicanery. Each 8k-token window contains 250 sampled positions, one probe every 32 tokens. The percentage is the fraction of those probes where the other backends highest-scoring token differed from Triton’s. Random noise was accounted for by running the same test with the same attention backend multiple times. The logits across runs at every hidden state were bit for bit identical. Meaning this particular divergence comes exclusively from the matrix multiplication and addition operations happening during prefill inside trt/fa2/fi. Disagreements appeared in clusters and varied with prompt content rather than increasing smoothly with context length. This is not evidence of one universal length at which the model “falls apart” but… we will get there soon… Now that we have a baseline comparison of interesting prompt fuel, lets dive into… Test 2: KV Cache quantization, or why your LLM’s IQ drops like a rock after 40k tokens Repeating the same methodology, we took the BF16 weights and BF16 kv cache baseline above running Triton, and ran the next experiment. What happens when you leave the weights and activations alone, and JUST quantize the kv-cache? Ah, divergence. And this leads us to our first dumpster-fire of the evening: a completely reproducible tool calling error. Enough top-tokens got flipped during tool calls, we let them play out and while BF16 was fine, int8 kv-cache eventually managed to recover, int4 did not! Test 3: Weight Weight, Don’t Tell Me! This time we are leaving all the kv-caches full size at bf16. We are adding some new players to the game however by comparing: - BF16 reference: Qwen/Qwen3.6-27B ( Qwen/Qwen3.6-27B · Hugging Face ) - Official FP8: Qwen/Qwen3.6-27B-FP8 ( Qwen/Qwen3.6-27B-FP8 · Hugging Face ) - INT8 W8A16: TheHouseOfTheDude/Qwen3.6-27B-INT8 ( TheHouseOfTheDude/Qwen3.6-27B-INT8 · Hugging Face ) - NVIDIA NVFP4: nvidia/Qwen3.6-27B-NVFP4 ( nvidia/Qwen3.6-27B-NVFP4 · Hugging Face ) - AWQ W4A16: cyankiwi/Qwen3.6-27B-AWQ-BF16-INT4 ( cyankiwi/Qwen3.6-27B-AWQ-BF16-INT4 · Hugging Face ) These 4 quants represent a broad picture of weights and activations. A notable piece of information for our mathnasium is the actual CUDA kernel / GEMM (general matrix multiplication) / MMA (matrix multiply accumulate) instructions being run to calculate the logits for each quant are different: Qwen3.6-27B (reference) - Weights/activations: BF16 weights, BF16 activations - Linear/GEMM: UnquantizedLinearMethod → torch.nn.functional.linear. Each CUDA tile selected by its associated shape/geometry. - KV cache: BF16 (Forced) - Qualification: Reference checkpoint. Qwen3.6-27B-FP8 - Weights/activations: E4M3 FP8 weights in 128×128 blocks; dynamic FP8 activation quantization inside converted linears; excluded modules such as lm_head remain BF16 - Linear/GEMM: Fp8LinearMethod → CutlassFp8BlockScaledMMKernel - KV cache: BF16 (Forced) - Qualification: DeepGemm was automatically disabled because vLLM flags its E8M0 scale format as accuracy-degrading for this architecture (SM120); CUTLASS was selected instead. No calibration dataset was identified in the published files. Qwen3.6-27B-INT8 - Weights/activations: Static, symmetric, channel-wise INT8 linear weights; BF16 activations (W8A16). GDN/linear_attn projections and lm_head excluded from quantization. - Linear/GEMM: CompressedTensorsWNA16 → MarlinLinearKernel - KV cache: BF16 (Forced) - Qualification: One-shot quantization with explicitly no calibration dataset. Its unusually good fidelity is less mysterious once you account for W8A16 plus unquantized GDN projections. Qwen3.6-27B-NVFP4 - Weights/activations: Mixed checkpoint — 208 static FP8 W8A8 targets covering 64 full-attention projections and 144 GDN projections; 193 NVFP4 W4A16 targets covering 192 MLP projections plus lm_head, group size 16 - Linear/GEMM: - FP8 targets: ModelOptFp8LinearMethod → FlashInferFP8ScaledMMLinearKernel - NVFP4 targets: NVFP4 GEMM → MarlinNvFp4LinearKernel - KV cache: BF16 (Forced) - Qualification: Not native FP4 arithmetic in our upstream-nightly run. vLLM classified the GPU path as lacking native FP4 support and explicitly selected weight-only FP4 compression through Marlin. The checkpoint’s embedded FP8 KV scheme was overridden with BF16 KV for the bakeoff. Qwen3.6-27B-AWQ-BF16-INT4 - Weights/activations: Static asymmetric INT4 weights, group size 32, MSE observer; BF16 activations (W4A16). GDN/linear_attn projections and lm_head excluded. - Linear/GEMM: CompressedTensorsWNA16 → MarlinLinearKernel - KV cache: BF16 (Forced) - Qualification: AWQ calibration dataset disclosed as “STEM and Agentic.” Other notable information for this run: - Full softmax/GQA attention for all models was AttentionBackendEnum.TRITON_ATTN; JIT monitor observed kernel_unified_attention. - GDN prefill: Triton/FLA GDN prefill kernel, requested as triton, head_k_dim=128. - During execution, the recurrent path also JIT-compiled _causal_conv1d_update_kernel, fused_recurrent_gated_delta_rule_packed_decode_kernel, and reduce_segments. - TP1, eager mode, no CUDA graphs, no MTP/speculative decoding, language-only execution. The next-token flip results shake out fairly predictably. TheDude (W8A16) mops the floor with everybody, beating first party FP8 (W8A8) and Nvidia(FP4-is-a-Lie) release. In fact, out of the 5 options, Nvidia’s release comes in dead last hitting ~50% token flips by the time we reach 88k context. Both the NVFP4 and AWQ W4A16 failed to properly close their tool calls and botched Cisco command line syntax (the correct command was ‘show arp’, while they executed ‘show run’), while both FP8 and INT8 were able to complete the correct calls. In future experiments I will try to explore the impact of using different fused GEMMs for the same weights, this is another interesting source of divergence where sometimes you have to trade precision for speed. Part 1 Wrap Up I have quite a few more experiments and observations to post, but require a great deal of parallel GPU time to calculate and record every logit sampled across huge context chains on multiple prompts with dozens of different settings. If you have specific questions, shoot me a DM or poke me on discord I guess.

2

JIT Compiling Code in 5μs

Hacker News · original → · 7/10 · AI/work: JIT compilation with AI for assembly generation
Historically, fast JIT compilation was a black art. To write a fast JIT compiler, you would need to know how to write assembly. Case in point: there is no production-ready database today that has…

Historically, fast JIT compilation was a black art. To write a fast JIT compiler, you would need to know how to write assembly. Case in point: there is no production-ready database today that has its own JIT compiler. They all either use LLVM or generate C/C++ code. Both of these options suffer from high compile times, which limits their applicability. Now, with the use of AI, it’s easier than ever to write a JIT compiler with fast compile times by directly targeting assembly. This is also one area of opportunity for new databases to improve on old ones. When building pgrust, I initially thought it would be really hard to implement a JIT compiler. In the end, I found it much easier than I expected due to AI assistance and it ends up being part of the reason why pgrust is so fast. The pgrust JIT compiler compiles code in around 5μs, which enables us to JIT compile every SQL query, not just a subset of them. In this post, I’ll walk you through how you can build your own fast JIT compiler. We’ll build a simple regular expression engine that uses JIT compilation as an example. Why JIT Compilation JIT compilation is the practice of generating compiled code at runtime or “Just In Time”. When done right, it can result in big performance wins, often on the order of 2-5x and sometimes even more. The main use case for JIT compilation is when there’s information you gain at runtime that drastically alters the behavior of your program. This is particularly common with programming language interpreters; they receive the code to execute at runtime. JIT compilers are also useful in domains beyond programming languages, such as parsing data. Sometimes you don’t know the schema of the data you’re parsing until runtime, and a JIT can help with that. To kick things off, let’s implement a toy regular expression engine. To keep things simple, we’ll support only two features: literal strings and repetition (i.e. the regex *). We’ll also skip the parser and represent the regular expression as already parsed Rust structures. This means we’ll be able to support strings such as: - apples - b(an)* but no alternation or lookbehind or anything like that. In code this is pretty simple. We’ll have 3 types of Nodes: a literal string node, a repetition node, and a concatenation node, which is the combination of two nodes. This ends up looking like this: enum Node { Literal(&'static str), Concatenation(Box<Node>, Box<Node>), Repetition(Box<Node>), } fn literal(text: &'static str) -> Node { Node::Literal(text) } fn concatenation(left: Node, right: Node) -> Node { Node::Concatenation(Box::new(left), Box::new(right)) } fn repetition(body: Node) -> Node { Node::Repetition(Box::new(body)) } Writing an interpreter for our regular expression engine is also straightforward: fn match_node(node: &Node, input: &[u8], pos: usize, next: &dyn Fn(usize) -> bool) -> bool { match node { Node::Literal(text) => { let literal = text.as_bytes(); input[pos..].starts_with(literal) && next(pos + literal.len()) } Node::Concatenation(left, right) => { match_node(left, input, pos, &|left_end| { match_node(right, input, left_end, next) }) } Node::Repetition(body) => { match_node(body, input, pos, &|body_end| { match_node(node, input, body_end, next) }) || next(pos) } } } fn interp_match(regex: &Node, input: &str) -> bool { let bytes = input.as_bytes(); match_node(regex, bytes, 0, &|pos| pos == bytes.len()) } Now this regular expression engine is pretty simple. It’s under 20 lines of code, but let’s see how it does in terms of performance. For comparison, we’ll compare the code against handwritten code implemented specifically for the regex. For our example we’ll use the regex b(an)*. The handwritten code ends up looking like: fn handwritten_b_an_star(input: &str) -> bool { let bytes = input.as_bytes(); let mut pos = 0; if pos == bytes.len() || bytes[pos] != b'b' { return false; } pos += 1; while pos < bytes.len() { if bytes[pos] != b'a' { return false; } pos += 1; if pos == bytes.len() || bytes[pos] != b'n' { return false; } pos += 1; } true } (There are ways you could optimize this code and make it much faster, but for our purposes it serves as a good comparison) When I benchmark a couple of examples against these two, I get that the handwritten version is 10-20x faster than the interpreter. Clearly a lot of room for improvement. Now let’s take a look at how we can use JIT compilation to get a general regular expression engine that performs as well as the handwritten version. How to JIT Compile There are two steps to JIT compile code. First you generate the assembly for the code you want to run. Once you have the code, you then package the assembly code into a function that you can call like any other code into your program. To generate the assembly, we will use a variant of an approach called copy-and-patch. The idea is that we have a series of templates in assembly for the different operations we want to JIT compile. These templates are called “stencils”. When we want to JIT compile an operation, we take the associated stencil and make small tweaks based on the specifics of the operation. Very similar to filling in a real stencil. By stringing together several of these filled stencils, we can construct a program at runtime that has similar performance to the handwritten version. Here’s the path we’ll take: first we’ll look at the ARM64 code we want to generate for b(an)*. Then we’ll turn repeated instruction sequences into reusable stencils, write an emitter that fills and combines those stencils from the regex AST, and finally copy the generated instructions into executable memory so Rust can call them like a normal function. To walk you through how this works, it’s easiest to start with the generated code and work backwards to the JIT compiler itself. Again, we’re working with the regex “b(an)*”. To lay out some design decisions: - We’ll use a stack for backtracking. The stack will keep track of the state we should go to if we hit a dead end in the regex - The string we are matching with will end in a null byte. That means any of our character comparisons will automatically fail if we hit the end of the string. This means we don’t have to do any length comparisons at any point For the state of our program we will use the following registers: - x0 – current position in string and return value - x1 – top of stack used for backtracking - x2 – bottom of stack used for backtracking (this is needed to determine if the stack is empty) - x9 – used as a temporary variable For the inputs into our program, we will be passed: - x0 – a pointer to the start of the string - x1 – a pointer to the location we will use for our stack Generated ARM64 Now that we’ve taken care of that, let’s walk through the generated assembly part by part. This is specifically on macOS with ARM64. First up, we have the prologue, which initializes the program. All it does is initialize the stack by setting the top of the stack and the bottom of the stack to the value passed in: 0: aa0103e2 mov x2, x1 Next up, we have the code that checks for the character b. If it sees a character that’s not b, we jump to a block of code that handles fallback logic. Otherwise, we advance our position in the string: ; CHAR 'b' 4: 39400009 ldrb w9, [x0] ; load current input byte 8: 7101893f cmp w9, #0x62 ; is it 'b'? c: 54000281 b.ne 0x5c ; no -> fallback block 10: 91000400 add x0, x0, #1 ; yes -> advance input Next up, we have the repetition (an)*. For the repetition, we need to do the backtracking. If we backtrack here, that means we jump immediately to the end of the loop. That means we need to store both the address of the instruction after the loop and our position in the string on the stack. 14: d2800989 movz x9, #0x004c ; build resume address 18: f2a00009 movk x9, #0x0000, lsl #16 ; = 0x1_0000_004c 1c: f2c00029 movk x9, #0x0001, lsl #32 ; (the loop exit) 20: f2e00009 movk x9, #0x0000, lsl #48 ; 24: a8810029 stp x9, x0, [x1], #16 ; push (exit, pos) onto stack With that in place, we can now execute the body of the repetition. This will check for the characters ‘a’ and ‘n’ and, if it sees them, go back to the top of the repetition, but at a new string location. ; CHAR 'a' 28: 39400009 ldrb w9, [x0] 2c: 7101853f cmp w9, #0x61 ; 'a'? 30: 54000161 b.ne 0x5c ; no -> fallback block 34: 91000400 add x0, x0, #1 ; CHAR 'n' 38: 39400009 ldrb w9, [x0] 3c: 7101b93f cmp w9, #0x6e ; 'n'? 40: 540000e1 b.ne 0x5c ; no -> fallback block 44: 91000400 add x0, x0, #1 ; JMP 48: 17fffff3 b 0x14 ; back to top of loop Now we’re past the loop. This is where the backtracking will jump once we backtrack. Once we finish the repetition, we’re at the end of the regex. All we have to do now is check if we’re at the end of the string. If we are at the end of the string, we return 1 for success. If we are not, that means the regex failed to match, and we need to run the fail logic to do a fallback. 4c: 39400009 ldrb w9, [x0] 50: 35000069 cbnz w9, 0x5c ; not at NUL -> fallback block 54: d2800020 mov x0, #1 ; success 58: d65f03c0 ret And then finally, we have the fallback logic. This checks if the stack is empty. If it is, we return 0. If it’s not empty, we pop both the fallback address and the fallback string position off the stack, and then jump to the fallback address. 5c: eb02003f cmp x1, x2 ; any frames left? 60: 54000060 b.eq 0x6c ; no -> give up 64: a9ff0029 ldp x9, x0, [x1, #-16]! ; pop (resume, pos) 68: d61f0120 br x9 ; jump there 6c: d2800000 mov x0, #0 ; no match 70: d65f03c0 ret Building the Stencils Now that you’ve had the chance to see the compiled code, you should start to get a sense of how the copy-and-patch compiler would work. We have common sets of instructions with only minor differences between them. For each of these blocks of functions, we can write a function to generate the respective code. Each function will take in values to use to modify the code. For example, one of the arguments to stencil_char will be the char in the regex to compare against. We’ll insert that char directly into the machine code. The prologue is straightforward since it’s just a block of code: const PROLOGUE_WORDS: usize = 1; fn stencil_prologue() -> [u32; PROLOGUE_WORDS] { [0xAA0103E2] // mov x2, x1 } For character comparison, we need to insert the character we’re comparing against and where to jump for the fallback logic: const CHAR_WORDS: usize = 4; fn stencil_char(byte: u8, stencil_pos: usize, fail_pos: usize) -> [u32; CHAR_WORDS] { [ 0x39400009, // ldrb w9, [x0] 0x7100013F | ((byte as u32) << 10), // cmp w9, #byte 0x54000001 | cond_branch_offset(stencil_pos + 2, fail_pos), // b.ne fail 0x91000400, // add x0, x0, #1 ] } For the repetition, we have the start of the loop that pushes onto the stack and the jump onto the end: const SPLIT_WORDS: usize = 5; fn stencil_split(resume_addr: u64) -> [u32; SPLIT_WORDS] { [ 0xD2800009 | addr_bits(resume_addr, 0), // movz x9, #addr[0..16] 0xF2A00009 | addr_bits(resume_addr, 1), // movk x9, #addr[16..32], lsl 16 0xF2C00009 | addr_bits(resume_addr, 2), // movk x9, #addr[32..48], lsl 32 0xF2E00009 | addr_bits(resume_addr, 3), // movk x9, #addr[48..64], lsl 48 0xA8810029, // stp x9, x0, [x1], #16 ] } const JMP_WORDS: usize = 1; fn stencil_jmp(stencil_pos: usize, target_pos: usize) -> [u32; JMP_WORDS] { [0x14000000 | branch_offset(stencil_pos, target_pos)] // b target } And then we have the match and fail blocks which are pretty clean: const MATCH_WORDS: usize = 4; fn stencil_match(stencil_pos: usize, fail_pos: usize) -> [u32; MATCH_WORDS] { [ 0x39400009, // ldrb w9, [x0] 0x35000009 | cond_branch_offset(stencil_pos + 1, fail_pos), // cbnz w9, fail 0xD2800020, // mov x0, #1 0xD65F03C0, // ret ] } const FAIL_WORDS: usize = 6; fn stencil_fail() -> [u32; FAIL_WORDS] { [ 0xEB02003F, // cmp x1, x2 0x54000060, // b.eq +3 (to the mov below) 0xA9FF0029, // ldp x9, x0, [x1, #-16]! 0xD61F0120, // br x9 0xD2800000, // mov x0, #0 0xD65F03C0, // ret ] } For completeness, here’s the helper functions we used which just help us insert specific data into the instructions: // Compute the branch-offset field for a conditional branch (b.ne / cbnz): // the instruction count from branch to target, stored in bits 5..24. fn cond_branch_offset(branch_pos: usize, target_pos: usize) -> u32 { let instr_count = target_pos as i64 - branch_pos as i64; // may be negative (((instr_count as u64) & 0x7FFFF) << 5) as u32 } // Compute the branch-offset field for an unconditional branch (b): // same idea, but stored in bits 0..26. fn branch_offset(branch_pos: usize, target_pos: usize) -> u32 { let instr_count = target_pos as i64 - branch_pos as i64; // may be negative ((instr_count as u64) & 0x3FF_FFFF) as u32 } // Extract 16 bits of an absolute address, positioned for a movz/movk immediate. fn addr_bits(addr: u64, part: usize) -> u32 { (((addr >> (16 * part)) & 0xFFFF) as u32) << 5 } Emitting Code Now the code that drives it: // Computes how many instructions a node compiles to. fn node_words(node: &Node) -> usize { match node { Node::Literal(text) => text.len() * CHAR_WORDS, Node::Concatenation(left, right) => node_words(left) + node_words(right), Node::Repetition(body) => SPLIT_WORDS + node_words(body) + JMP_WORDS, } } struct Emitter { code: Vec<u32>, fail: usize, // word offset of the shared fail block base: u64, // runtime address of code[0], for absolute-address holes } impl Emitter { // Returns the offset where the next instruction will be placed. fn pos(&self) -> usize { self.code.len() } // Appends a filled stencil to the code buffer. fn emit(&mut self, stencil: &[u32]) { self.code.extend_from_slice(stencil); } // Emits the code for one node, recursing into children. fn emit_node(&mut self, node: &Node) { match node { Node::Literal(text) => { for &byte in text.as_bytes() { self.emit(&stencil_char(byte, self.pos(), self.fail)); } } Node::Concatenation(left, right) => { self.emit_node(left); self.emit_node(right); } Node::Repetition(body) => { let split_at = self.pos(); let exit = split_at + SPLIT_WORDS + node_words(body) + JMP_WORDS; self.emit(&stencil_split(self.base + exit as u64 * 4)); self.emit_node(body); self.emit(&stencil_jmp(self.pos(), split_at)); } } } } // Generates the complete program: prologue, the compiled AST, MATCH, fail block. fn generate_code(regex: &Node, base: u64) -> Vec<u32> { let nwords = PROLOGUE_WORDS + node_words(regex) + MATCH_WORDS + FAIL_WORDS; let mut emitter = Emitter { code: Vec::with_capacity(nwords), fail: nwords - FAIL_WORDS, base, }; emitter.emit(&stencil_prologue()); emitter.emit_node(regex); let match_at = emitter.pos(); emitter.emit(&stencil_match(match_at, emitter.fail)); emitter.emit(&stencil_fail()); assert_eq!(emitter.pos(), nwords); emitter.code } And that’s the hard part! Personally, writing assembly is where I find AI the most helpful. My main experience with assembly is completing the microcorruption CTF. I’ve never actually written assembly myself. I would really struggle to figure out the exact instructions needed and how to modify them to get the output I wanted. With AI, I can give my coding agent the general shape of how I want the JIT compiler to work, and it can handle a lot of these details for me. Loading Machine Code To finish our compiler we need to actually load the code. To do this, we’ll use mmap to allocate a block of memory that is readable, writable, and executable. We’ll then copy the code into that memory and convert that block of memory into a function which we then call: const BSTACK_MAX: usize = 4096; // These functions are included in the mac system library unsafe extern "C" { fn pthread_jit_write_protect_np(enabled: libc::c_int); fn sys_icache_invalidate(start: *mut libc::c_void, len: libc::size_t); } type MatchFn = unsafe extern "C" fn(input: *const u8, bstack: *mut u64) -> u64; struct Jit { buf: *mut u32, nbytes: usize, bstack: Vec<u64>, } impl Jit { fn compile(regex: &Node) -> Jit { let nwords = PROLOGUE_WORDS + node_words(regex) + MATCH_WORDS + FAIL_WORDS; let nbytes = nwords * 4; unsafe { let buf = libc::mmap( std::ptr::null_mut(), nbytes, libc::PROT_READ | libc::PROT_WRITE | libc::PROT_EXEC, libc::MAP_PRIVATE | libc::MAP_ANON | libc::MAP_JIT, -1, 0, ) as *mut u32; assert!(buf as *mut libc::c_void != libc::MAP_FAILED, "mmap failed"); let code = generate_code(regex, buf as u64); pthread_jit_write_protect_np(0); // make the region writable (Apple W^X) std::slice::from_raw_parts_mut(buf, code.len()).copy_from_slice(&code); pthread_jit_write_protect_np(1); // back to executable sys_icache_invalidate(buf as *mut libc::c_void, nbytes); Jit { buf, nbytes, bstack: vec![0; BSTACK_MAX * 2] } } } // Runs the generated code. Input must end with a NUL byte. fn is_match(&mut self, nul_terminated: &[u8]) -> bool { debug_assert_eq!(nul_terminated.last(), Some(&0)); unsafe { let matcher: MatchFn = std::mem::transmute(self.buf); matcher(nul_terminated.as_ptr(), self.bstack.as_mut_ptr()) != 0 } } } impl Drop for Jit { fn drop(&mut self) { unsafe { libc::munmap(self.buf as *mut libc::c_void, self.nbytes); } } } Results With all of this complete, let’s compare the performance of the different implementations we built: | Input length | Interpreter | JIT | Handwritten | JIT speedup | Handwritten speedup | |---|---|---|---|---|---| | 9 | 45 ns | 3.8 ns | 3.8 ns | 11.7x | 11.9x | | 33 | 103 ns | 7.9 ns | 10.5 ns | 13.0x | 9.8x | | 129 | 597 ns | 30 ns | 32 ns | 19.7x | 18.6x | | 513 | 1,955 ns | 126 ns | 120 ns | 15.5x | 16.2x | | 2,049 | 8,301 ns | 470 ns | 393 ns | 17.7x | 21.1x | So JIT and the hand-rolled implementation are pretty much neck and neck. Sometimes the JIT version is faster, and sometimes the hand-rolled version is faster. There’s been a meme circulating about how AI doesn’t help because “code was never the hard part.” I think that’s true in some domains, but in others, writing the code absolutely was the hard part. JIT compilers are a great example of that. For many pieces of software, a JIT compiler would help a lot with speeding up the code. The rarity of JIT compilers makes me believe that implementing a JIT compiler historically was too difficult for it to be worthwhile. LLMs have lowered the barrier to entry and made it much easier to write a JIT compiler. This is the thesis behind pgrust. Databases historically were the hardest piece of software to build and were limited because of that. Now, with AI, we can be more ambitious about the type of software we build. Thanks for reading, and if you want to support the project, the best way to support pgrust is to give us a star on GitHub. If you want to follow along:

3

I set a trap for a book-marketing scammer (2025)

Hacker News · original → · 7/10 · AI: scam detection using AI agents and identity theft
This is the story of a traditionally published sci-fi author who received more than fifty scam book-marketing pitches in thirty-three days, all hoping to prey on his desperation to sell books and…

This is the story of a traditionally published sci-fi author who received more than fifty scam book-marketing pitches in thirty-three days, all hoping to prey on his desperation to sell books and net his next publishing deal. We’ll call him Rob. It’s also the story of a plucky, persistent, puzzled AI named Veronica and a scammer who steals real authors’ identities in order to get their marks’ guards down. But mostly, it’s the story of what happens when an entire industry is drowning, and the people holding out life preservers are running a con. Day Zero: Oct. 18 The wave began with Mercy Gold (which is a great name for a werewolf hunter or demon slayer, BTW). “Hello,” her email started. “You’ve already laid a great foundation with your blurb and keywords that’s a strong start.” A compliment sandwich, I know them well. You’re doing great, BUT there’s a problem you didn’t know you had. She offered to help me manage “light social presence and simple email marketing” to keep my book “in front of real readers.” Would I be open to seeing how this could support my “author journey”? The email came from mercygold7400@gmail.com. Not a company domain. Not a professional address. Just a numbered Gmail account. I deleted it. Two days later, another pitch arrived. Then another. By the end of October, I’d received seven. In the first three weeks of November, I got twenty-three more. They came from: Isaac Michael, offering Pinterest and Goodreads strategies Halfdaytravel, promising TikTok features at 1:28 in the morning Pratibha Malav, with a price list for paid reviews across multiple platforms Janet, pitching websites and cinematic trailers Someone claiming to be “Dr. Sandra Maria Anderson” using the email address abbyamazonalirah@gmail.com Some offered video production, others social media management, still others Amazon optimization or Google indexing services. They all shared generic Gmail addresses, vague promises of increased visibility, and an understanding that I was anxious about my books finding readers. They were 100-percent right about the anxiety, and that’s what makes them dangerous. The economics of traditional publishing are brutal right now, especially for mid-list (and lower) writers. My actual royalties per book sold (assuming they’ve earned out and I get, ya know, actual royalties): Ebook: roughly $0.25-0.40 Print book: roughly $0.65-1.00 Audiobook: roughly $1.30-5.00 (depending on format and territory) If someone charges me $500 for marketing services, I’d need to sell as many as 2,000 additional copies just to break even. Publishers have PR teams, media contacts, co-op placement deals with retailers, and marketing budgets. If Angry Robot, with all those resources, couldn’t get my book onto certain Pinterest boards or Goodreads lists, what makes a freelancer with a Gmail account believe their strategy will work? The truth is, they don’t believe it will. They believe I might believe it will. Books ARE harder to discover than they used to be. Algorithms DO matter. Readers ARE out there, somewhere, not finding my work. And, when the next pitch arrives—probably tomorrow or right now—some part of me will wonder: what if? Glub. Glub-Glub. Glub-Glub-Glub. To understand why these scams are proliferating, look at what’s happening to the industry. More books are being published than ever before (self-publishing explosion plus traditional publishers desperate for revenue), while readership is declining (fewer readers, less reading time, more competition from other media). The result is the average book sells fewer copies than it used to. Authors are desperate because books aren’t selling, traditional marketing isn’t working, and publishers are offering less support per title. We’re feeling pretty helpless about our career trajectories, watching sales numbers decline while wondering what we’re doing wrong. Meanwhile, publishers are desperate because the mid-list is dying, there are fewer breakouts, and they can’t afford to market every title the way they used to. They tell authors there’s nothing more they can do, leaving us to wonder if we should be doing something ourselves. Scammers see opportunity in this desperation. There are more anxious authors than ever before, each one a potential customer. The market is growing (more books = more targets) while competition for attention intensifies. It’s the perfect environment to scale up operations using AI and automation. That’s why the Summer and Fall of 2025 saw an explosion of these pitches. Summer sales numbers were bad across the industry. Publishers quietly told authors they weren’t picking up next options. Industry articles about the “publishing crisis” circulated widely. The anxiety peaked just as the Q4 season began, and the scammers launched coordinated campaigns targeting authors most likely to be feeling desperate. In recent articles, Victoria Strauss, co-founder of Writer Beware (the industry watchdog that protects authors from scams, sponsored by Science Fiction & Fantasy Writers Association), has confirmed what I was experiencing wasn’t isolated. She’s been tracking the same wave since June 2025, tracing it to operations in Nigeria using AI to generate personalized pitches at scale. The scams, she wrote, had ramped up faster than any fraud in her decades of experience. Nov. 3: The Flood Five, count-em, five book marketing pitches in my inbox. One had arrived at 1:22am while I slept, from someone called “Green Link” promising to feature my book in “active reading communities” through “real organic [Hence the Green? Or am I the Greene?] no ads, just real engagement.” Another came at 9:48am from Sylvester Josh, offering 3D cinematic book trailers. By afternoon, I’d been pitched TikTok features, paid reviews that violated Amazon’s policies, and Amazon optimization services for problems I didn’t have. A sixth pitch arrived that day but was caught by Gmail’s spam filter—someone claiming to be “Dr. Sandra Maria Anderson” using an email address that didn’t match the name at all. Nov. 3 was a preview. Over the next three weeks, the pitches would arrive almost daily, sometimes multiple times per day, from senders with names like Mercy Goldcrown (when plain Gold isn’t enough), Naphy Expertt, Ophelia, Bella, Emmanuel, Evelyn, and KAMBIO. They offered every service imaginable. The prices ranged from $20 to “contact me for [a] quote.” They all claimed to have “real readers,” “active communities,” and “organic engagement.” None provided portfolios, case studies, or verifiable results. Gmail’s SPAM filter had been fighting for me, catching roughly two out of every three pitches. But that still meant dozens were getting through. Between October 18 and November 20—just thirty-three days—I documented at least fifty-one pitches. The real number was probably higher. Some I’d deleted without recording. Some had been auto-deleted from spam after thirty days. Some I’d forgotten about before I started keeping systematic track. I wasn’t being marketed to. I was being carpet-bombed. A Note on Self-Published Authors Self-published authors face an even more intense version of this siege. With higher per-book royalties (often $2-7 per ebook versus my $0.25-0.40) and direct control over their Amazon pages, they’re more tempting targets—and the math looks just plausible enough to be dangerous. “If I make $3 per book and sell 200 extra copies, that’s $600!” But they still need to sell those 200 ADDITIONAL copies directly attributable to the service, and fake Pinterest boards don’t work any better for self-pub than they do for traditional publishing. Self-pub authors are also more vulnerable because they’re responsible for everything: cover design, formatting, distribution, marketing. When someone offers to handle “just the marketing part,” it’s tempting. They have no publisher to shield them, no agent to offer advice, no industry contacts to warn them away. Victoria Strauss (Writer Beware) has noted that earlier waves of publishing scams—from the Philippines and Pakistan—focused almost exclusively on self-published authors. The Nigerian operations are different: they target everyone with a published book, regardless of how it got there. It’s a Trap! (Mine) On Oct. 28, someone named Isaac Michael sent me seven messages in a single day, each one arriving within hours despite my repeated, polite “no thank you” responses. He’d promised to optimize my “visibility window”—I had exactly sixty days before the algorithm reset, he claimed. He’d already mapped the strategy: 12 Pinterest boards, 8 Goodreads lists, full sequencing. He cited specific boards: “Retro Sci-Fi Worlds” with 18K+ saves, “Alien First Contact Stories” with 12K+ saves. His response speed suggested AI assistance. The pressure tactics suggested desperation. The numbered Gmail account (isaacmichael0181@gmail.com) suggested the same operation I was seeing everywhere else. (He also tried to convince me to spend all this money on the second book of a duology, which made no fucking sense.) I finally ignored Isaac. But on Nov. 13, the same playbook returned with a new sender. “Hey Robert,” Veronica Emmanuel began, “I’ve been digging into Six Plays and noticed a couple of big missed opportunities.” She claimed the book was missing opportunities on Pinterest and Goodreads. She’d identified specific boards and lists where the work should appear. She had the complete strategy ready: sequencing, timing, a plan to sync both algorithms within a critical sixty-day window. It was detailed and professional-sounding. The same service Isaac had pitched, delivered by someone new! Except for one problem: I have never written a book called Six Plays. Six Plays is a collection of works by Robert Greene—the Elizabethan playwright who died in 1592. Veronica Emmanuel, with her detailed Pinterest strategies and algorithm insights and carefully mapped visibility plans, had confused me with a man who’d been dead for 433 years. Seven days later, she followed up. Had I reviewed her strategy for Six Plays? The timing was critical. The algorithm window was closing. I decided to test the system. “Hi Veronica,” I replied. “I’m very interested in your strategy for Six Plays. Just one question: which edition are you working with? The 1599 quarto or the 1861 Dyce compilation? Also, do you think Pinterest is the right platform for Renaissance drama, or should we focus on the Globe Theatre’s social media? Thanks, Robert Greene (the science fiction author, not the dead playwright).” Her response arrived within hours. She’d chosen the 1861 Dyce compilation, she explained, for its “cohesive editorial structure” that made it easier to build marketing materials. Pinterest wasn’t conventional for Renaissance drama, but it worked through “visual discovery and mood-based storytelling.” She cited specific boards: “Renaissance Theatre Moodboards” with 8-12K monthly saves, “Historical Drama Visuals” with 10K saves, “Shakespeare-Inspired Imagery” with 7-9K saves. The response was sophisticated, detailed, confident. It explained audience overlap with “dark academia” aesthetics, compared Pinterest’s algorithmic distribution to the Globe Theatre’s narrower reach, and discussed how to bridge “literary integrity and modern reader discovery.” It was impressive. It was also completely insane. “Thanks for your questions,” she’d written, “and noted, the sci-fi author, not the Renaissance playwright!” Then she’d spent three paragraphs explaining her marketing strategy for the Renaissance play collection anyway. I gave her one more chance. “Veronica, I have never written a book called Six Plays. That’s a collection by an Elizabethan playwright who died in 1592. I yet live and I write science fiction. Can you explain what’s happening here?” Two and a half days later, she apologized. There had been a mixup, she explained. She worked with “a large number of authors each week” in batches, and my name had come through in a group where someone was requesting analysis of the public-domain Six Plays collection. “The system flagged your name as connected to that title,” and she’d mistakenly continued the conversation. The system. She claimed to be “familiar with my traditionally published novels,” though she never named them. She offered to “regroup and talk about my actual catalog.” The sale was still on, if I was interested. But I’d learned what I needed to know. This wasn’t a person carefully researching authors and crafting custom strategies. This was an operation using AI to process names in batches, generating confident pitches, powered by systems that couldn’t tell the difference between a living science-fiction writer and a playwright who’d been dead for four centuries. And when caught in an impossible error, the response was: apologize, blame the system, pivot to the actual product. Veronica followed up twice more in the next twenty-four hours, offering generic “multi-platform frameworks” and “audience targeting strategies.” She still never named my books. She claimed to have identified “time-sensitive opportunities” in my catalog and put together a breakdown of where I was “missing discoverability.” The Identity Thief While I was fencing with Veronica, another scammer was trying a different approach: pretending to be someone real. On Nov. 2, I got an email from “Judy Leigh” writing from the address faithexpert92@gmail.com. “Hi, I’m Judy, an author based in the UK. I published my book not too long ago, and it has recently started performing very well and it’s currently one of the top-selling books on Amazon in its category.” She wanted to connect with fellow authors to share strategies. Would I be interested in learning how she was growing her readership and sales? I might have been flattered to hear from such a fancy person, but the email address was suspicious: “faithexpert92” had nothing to do with “Judy Leigh.” Then on November 18, “Judy Leigh” followed up. She was checking in to see if I’d received her first message about book marketing strategies. Would I like to connect? Something made me Google the name. There IS a Judy Leigh. She’s a legitimate UK author published by Boldwood Books who has sold over a million copies. She writes uplifting contemporary fiction celebrating friendship and second chances. She also writes historical novels under the pseudonym Elena Collins. Faithexpert92 was even using Judy Leigh’s headshot as a profile picture. But I was highly suspicious that the person emailing me wasn’t her. I decided to test this one too. I lied apologetically—sorry for the delay, your message went to spam. Then I asked a simple question: “How did you find out about me?” Four hours later, “Judy” responded with “Hi Greene, Thank you so much for getting back to me!” Instagram had suggested my page, she claimed. Since we were both authors, my work had caught her attention. She asked about my sales and what marketing strategies I’d implemented—standard sales qualification questions to assess whether I was desperate enough to buy. I played along, positioning myself as the perfect target (which I am): traditionally published, relying on my publisher for marketing, passive about promotion, seemingly naive. “Ha! Your email address threw me off a bit. I thought you might be a scammer. It’s nice to know you are a real person! My publisher does the lion’s share of the marketing chores. I just write the things!” The next morning, “Judy” delivered her pitch. She’d “recently discovered” that Medium was running an “End-of-Year Book Feature Promotion” selecting 200-300 books for automated daily promotion. The window was closing soon. She’d tried it herself and seen “a noticeable boost.” There was one detail: she worked with “someone in the United States who specializes in writing book-related blog posts on Medium.” They’d handled hers. Would I like the contact? This was the pattern Victoria Strauss had documented: scammer builds trust, offers helpful advice, then refers to a third party for payment. The money routes through Upwork or Fiverr to an “assistant” in Nigeria. The service either doesn’t materialize or delivers something worthless. I’d also seen this exact pitch before. Three days earlier, someone called “Mercy Goldcrown” had offered me a “$30 promo slot for a short, polished announcement blog” on Medium, claiming it was perfect for the “Ember season.” Medium doesn’t run “End-of-Year Book Feature Promotions.” The entire premise was fabricated. I messaged the real Judy Leigh through Facebook, then emailed her publisher Boldwood Books to alert them that someone might be using her identity to scam other authors. Neither responded. Finally, I posted a warning to the real Judy’s blog, explaining that someone was impersonating her using an email address with her headshot. She let the comment go through the moderation process and responded with a “Like.” No comment. No outrage. No warning to her readers. Just acknowledgement. (Did she know? Is she tired of hearing about it? I’ve no way to tell!) The fake Judy Leigh, meanwhile, was still waiting for my response about the “expert” who could help get my book selected for Medium’s non-existent promotion. On Nov. 25, I told her the truth: “’Judy’: Medium doesn’t run ‘End-of-Year Book Feature Promotions.’ I know because I checked. I also checked with the real Judy Leigh and her publisher to see if you were she. Outlook not so good. So, I’ve been documenting this entire exchange for an investigative article about book marketing scams. Would you like to comment?” She never responded. Addendum: Writer Ann Leckie, who I follow on BlueSky, recently reported someone was using her name to run a similar scam. Lies and Terrible Realities The cruelest part isn’t that the scammers are lying about their services. Scammers gotta scam. It’s that the problem they’re claiming to solve is real. Books ARE harder to discover. Algorithms DO matter. The gap between authors and readers IS widening. Traditional marketing IS failing. And the scammers are exploiting this with sophisticated psychological hooks: They identify real problems (discoverability, algorithm changes, market saturation) They offer specific solutions (Pinterest boards with exact save counts, Goodreads lists with member numbers) They create urgency (algorithm windows closing, seasonal opportunities, competitive selection processes) They minimize barriers (low prices, “no time required on your end,” pre-built strategies) They exploit isolation (authors working alone, desperate for answers, willing to try anything) These services can’t fix anything, but there’s always someone willing to risk money on a map to buried treasure. The Long, Dark Night of the Soul As I finished writing this last week, three more pitches arrived. One offered TikTok features. One promised Amazon visibility optimization. One claimed to run a “curated community of 5,000+ passionate readers” who will review my book for small tips of $25-30 each, minimum 30 reviewers required. Thirty times thirty is -- what? -- $900? I could foot that. Because what if this one is different? What if I’m wrong? What if shelling out a grand is all I need to get my writing career really going? What if I’m the fool for not even trying? Man, I want to believe I can change my stars with a simple cash infusion! That little inner whisper is what they’re counting on: that the gap between the work we’ve done and the readers who haven’t found it will always be wide enough for hope to slip through. That desperation can overwhelm mathematics. Tomorrow, another author will receive their first pitch. They’ll wonder if maybe, just maybe, this could be the thing that works. Somewhere, a system will flag their name, generate a personalized email, and send it from a numbered Gmail account. The algorithm never sleeps. The pitches never stop. Here’s one now. And now. Now. Note: If you’re an author receiving similar pitches, Victoria Strauss’s Writer Beware blog (writerbeware.blog) maintains updated information about current scams. The Science Fiction and Fantasy Writers Association (SFWA) also provides resources for identifying and reporting fraud. Most importantly: talk to other authors. The community is our best defense. Rob Greene (R.W.W. Greene) is the author of several science-fiction novels published by Angry Robot Books. His newsletter “twenty-first-century blues” examines culture, technology, art, and the systems that shape our lives. In writing this article, Greene used Claude (Anthropic) as a tool for structure, data organization, and feedback. I've received hundreds of those emails, too. Love to see what their AI comes up with - so many different versions of praise. I hadn't paid attention to the numbered gmail address, but now that you've pointed it out, I see that on the majority of ones I've saved. (I just have to save some!) I have received hundreds of emails just like the ones you described. Someone claiming to be Cal Newport has reached out to me on Bluesky, Tumblr, and my email using a numbered Gmail account. His publicist verified to me that it wasn't him. Writer Beware is a valuable resource. Thanks for sharing this.

4

🎙️ How I AI: How a solo founder used Codex and ChatGPT to launch a fashion brand without engineers

Lenny's Newsletter · original → · 7/10 · AI/SaaS: solo founder using AI to build fashion brand
[image →]How a solo founder used Codex and ChatGPT to launch a fashion brand without engineers | Yana WelinderListen now on YouTube • Spotify • Apple Podcasts[image →]Brought to you…

How a solo founder used Codex and ChatGPT to launch a fashion brand without engineers | Yana Welinder

Listen now on YouTube • Spotify • Apple Podcasts

Brought to you by:

  • Merge—Connective infrastructure for production AI

  • Jira AI SDLC—Get your tokens’ worth with Jira

Yana Welinder is the solo founder of Yana Bana, an AI-native fashion brand she’s building without an engineering team. In this episode, she walks through how she uses ChatGPT and Codex to turn hand-drawn sketches into realistic product images, operate professional 3D fashion software, research and contact manufacturers, and build a complete e-commerce site with voting and payments.

Biggest takeaways:

  1. The prompt is the spec, and the spec is everything. Before generating a single image, Yana creates a detailed “fashion prompt” describing the silhouette, the way the fabric should behave and move, and even the sound it should make. That upfront work leads to dramatically better results. The lesson applies far beyond fashion: the more clearly someone can define what “good” looks like, the better both AI and humans can deliver it.

  2. For designers, following the vision matters more than producing something impressive. Many image models can create a beautiful garment, but it often looks like something that already exists. ChatGPT Images 2.0 stays much closer to Yana’s original sketches. For creative professionals with a specific point of view, accurately following the design is far more valuable than generating something flashy.

  3. AI agents can make specialized software accessible without years of training. Yana used Codex to operate CLO, a professional 3D fashion design tool, and create the CAD files she needed for 3D printing, even though she had never learned to use the software herself. This points to a much bigger shift: people may no longer need to master every complicated tool before they can produce professional work with it.

  4. AI is making previously impractical ideas possible. Yana designed a Ruth Asawa–inspired gown featuring large sculptural forms that would have been extremely difficult to produce before. The limitation was not the lack of 3D printing but the enormous amount of CAD work required to prepare the design. AI removed that bottleneck. Similar bottlenecks are hiding in nearly every industry, often preventing great ideas from being worth pursuing.

  5. The near-term future is agents working through purpose-built software. Codex cannot create a finished CAD file on its own. But Codex operating CLO can produce exactly what Yana needs. That same pattern appears throughout her workflow: the AI serves as the orchestration layer, while specialized software handles the execution. SaaS is not disappearing; it is gaining a new kind of user.

  6. Voice-first, asynchronous work is already changing what one person can accomplish. Yana can start a deep research task, step away to sew or drape fabric, then return and continue the work through voice. The AI keeps working while she focuses on the physical parts of her business. For a solo founder, this kind of asynchronous collaboration can dramatically expand what fits into a single day.

  7. When the best path is unclear, humans and AI can work in parallel. One of Yana’s hardest remaining challenges is turning her designs into accurate sewing patterns. Instead of betting entirely on one approach, she has human patternmakers and Codex working on the problem at the same time. It is a practical way to move quickly under uncertainty: test both paths, compare the results, and let the better solution win.

Blog and detailed workflow walkthroughs from this episode:

Workflows for an AI-Native Fashion Brand: https://www.chatprd.ai/how-i-ai/workflows-for-an-ai-native-fashion-brand
↳ How to Use AI for Fashion Design and Visualization: https://www.chatprd.ai/how-i-ai/workflows/how-to-use-ai-for-fashion-design-and-visualization
↳ How to Build and Run an E-Commerce Business with AI as a Technical Co-Founder: https://www.chatprd.ai/how-i-ai/workflows/how-to-build-and-run-an-e-commerce-business-with-ai-as-a-technical-co-founder
↳ How to Prototype Complex Garments Using AI and 3D Modeling: https://www.chatprd.ai/how-i-ai/workflows/how-to-prototype-complex-garments-using-ai-and-3d-modeling


If you’re enjoying these episodes, reply and let me know what you’d love to learn more about: AI workflows, hiring, growth, product strategy—anything.

Catch you next week,
Lenny

P.S. Want every new episode delivered the moment it drops? Hit “Follow” on your favorite podcast app.

5

How a solo founder used Codex and ChatGPT to launch a fashion brand without engineers | Yana Welinder

Lenny's Newsletter · original → · 7/10 · AI/SaaS: AI-native fashion brand without engineers
Yana Welinder is the solo founder of Yana Bana, an AI-native fashion brand built with AI as her technical co-founder, starting from hand-drawn sketches and ending with runway photos, CAD files for…

Yana Welinder is the solo founder of Yana Bana, an AI-native fashion brand built with AI as her technical co-founder, starting from hand-drawn sketches and ending with runway photos, CAD files for 3D printing, and a live Stripe-connected pre-order site—no engineers required. A former product leader, she brings an operator’s rigor to her creative process: her “fashion prompt” is a detailed spec covering silhouette, volume, fabric beavior, movement, and sound, and watching her use Codex plus computer use to navigate 3D design software that’s entirely new to her is a clarifying demo of what today’s toolset actually makes possible.

Listen or watch on YouTube, Spotify, or Apple Podcasts

What you’ll learn:

  1. How Yana uses a custom fashion prompt as a technical spec to get consistent, realistic, on-design outputs

  2. Why ChatGPT Images 2.0 outperforms other models for fashion design

  3. How she uses Codex plus computer use to operate CAD and fashion software she’s never personally learned

  4. The workflow for taking a garment from hand-drawn sketch to product photo, runway photo, and influencer shot in a single session

  5. How she ran vendor outreach end to end using deep research and browser use

  6. How she built a full e-commerce site with voting, databases, and Stripe integration

  7. Why she’s testing human patternmakers and Codex in parallel


Brought to you by:

Merge—Connective infrastructure for production AI

Jira AI SDLC—Get your tokens’ worth with Jira

In this episode, we cover:

(00:00) Introducing Yana Welinder and Yana Bana

(02:38) Tour of the Yana Bana site

(05:20) The fashion prompt stack

(07:39) Live demo: generating a jacket from a prompt in ChatGPT

(10:01) Why Image Gen 2.0 beats other models

(11:51) The “prompt as spec” principle

(14:02) Iterating the design

(17:12) Using Codex and computer use to build CAD files in 3D software

(20:50) Vendor research, outreach emails, and Superhuman browser use

(23:34) Building the full e-commerce site

(27:40) Quick recap and what’s still hard

(30:05) How Yana prompts when AI pushes back

(31:15) Where to find Yana and how to vote on her garments

Tools referenced:

• ChatGPT (Images 2.0): https://chat.openai.com

• Codex (OpenAI): https://openai.com/codex

• CLO 3D (fashion pattern software): https://www.clo3d.com

• Vercel: https://vercel.com

• GitHub: https://github.com

• Stripe: https://stripe.com

• Superhuman: https://superhuman.com

Other references:

• Ruth Asawa: https://ruthasawa.com

• SFMOMA (Ruth Asawa): https://www.sfmoma.org/artist/Ruth_Asawa/

Where to find Yana Welinder:

LinkedIn: https://www.linkedin.com/in/ywelinder/

X: https://x.com/yanabana

Website: https://www.yanabana.com

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

6

SOURCES: Poolside Strikes $6 Billion Licensing Deal with Nvidia & Raises $1 Billion for Remaining Company at $12 Billion Valuation

Newcomer · original → · 7/10 · AI/platforms: Poolside AI licensing deal with Nvidia
[image →]EXCLUSIVE TO NEWCOMER: Poolside AI, the artificial intelligence model-building startup, has struck a non-exclusive licensing deal with Nvidia for $6 billion, plus a $1 billion investment in…

EXCLUSIVE TO NEWCOMER: Poolside AI, the artificial intelligence model-building startup, has struck a non-exclusive licensing deal with Nvidia for $6 billion, plus a $1 billion investment in Poolside at a $12 billion pre-money valuation, according to a letter to investors obtained by Newcomer.

Read more

7

The Story of a Cap Table: OpenRouter

Newcomer · original → · 7/10 · AI/platforms: OpenRouter acquisition by Stripe
[image →]OpenRouter’s founder Alex Atallah and AI investor Anjney Midha became fast friends over 13 years ago at Stanford, but it took more than a decade for them to partner up professionally — and…

OpenRouter’s founder Alex Atallah and AI investor Anjney Midha became fast friends over 13 years ago at Stanford, but it took more than a decade for them to partner up professionally — and the results have proven spectacular.

Stripe announced Wednesday its agreement to purchase the model routing platform, co-founded by Atallah in 2023. Midha had backed the team in 2025 while he was a general partner at Andreessen Horowitz. It’s the type of rapid-fire founding-to-multibillion-dollar-exit story the AI boom has supercharged in Silicon Valley.

We got the details on the deal and OpenRouter’s cap table from sources close to multiple firms with stakes in the company, all of whom requested anonymity to discuss sensitive financial information. A few key numbers:

Read more

8

Quoting Linus Torvalds

Simon Willison · original → · 7/10 · AI: Linus Torvalds using AI for kernel debugging
22nd August 2026 And this was a debug session from hell, enormously helped by an AI doing much of the grunt-work. I'd like to call it my tireless helper, but the AI several times stated flat out…

22nd August 2026 And this was a debug session from hell, enormously helped by an AI doing much of the grunt-work. I'd like to call it my tireless helper, but the AI several times stated flat out that this was impossible and unsolvable and that we should just write a report about it. I suspect those things have been trained by people who may not be quite as stubborn as I am. But while the AI was ready to give up several times, it did keep adding debug code and analyzing it faithfully when I pushed. So credit where credit is due and I let the AI write the commit message above. — Linus Torvalds, drm/xe: Don't hand out the flat CCS storage as usable VRAM Recent articles - Conceptual integrity and counting lines of code - 19th August 2026 - Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things - 16th August 2026 - Now we have a timeline of the OpenAI accidental attack against Hugging Face - 7th August 2026

9

llm 0.33

Simon Willison · original → · 7/10 · AI/platforms: LLM embedding model improvements
22nd August 2026 My highlights from this release: I shipped a quick 0.32.1 fix for this yesterday, but this is the more comprehensive fix. llm embed andllm embed-multi now accept--key . The…

22nd August 2026 My highlights from this release: I shipped a quick 0.32.1 fix for this yesterday, but this is the more comprehensive fix. llm embed andllm embed-multi now accept--key . The PythonEmbeddingModel.embed() ,EmbeddingModel.embed_multi() ,Collection.embed() andCollection.embed_multi() methods acceptkey= too, passing the resolved per-call key to embedding plugins without changing shared model state. Existing plugins that readself.key continue to work through a compatibility fallback. Thanks, ChrisJr404. #757, #1620 The embedding models now use the same pattern for keys that regular LLM models do. llm prompt -t/--template can now be repeated to combine templates in order. This allows model configuration and options from one template to be used with a prompt from another. This unlocks a neat pattern where you can create templates that package a model with a set of default options: llm -m gpt-5.6-luna -o reasoning_effort high --save lhigh llm "Generate an SVG of a pelican riding a bicycle" --save pelican # Combine and run the templates llm -t lhigh -t pelican - Reasoning-capable Responses API models now support a reasoning_summary option withauto ,concise , anddetailed values. This can be used with llm openai endpoint --responses. #1600 This is particularly useful for exercising different models that provide their own imitation of the OpenAI Responses API. Recent articles - Conceptual integrity and counting lines of code - 19th August 2026 - Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things - 16th August 2026 - Now we have a timeline of the OpenAI accidental attack against Hugging Face - 7th August 2026

10

More than just code review

Simon Willison · original → · 7/10 · AI/work: code review and AI agent verification practices
22nd August 2026 The key skill required to make productive use of coding agents is being able to confidently instruct them on how to make changes and then confidently verify that those changes have…

22nd August 2026 The key skill required to make productive use of coding agents is being able to confidently instruct them on how to make changes and then confidently verify that those changes have been applied in the correct way. Sometimes this involves reviewing every line of code they have written, but there are other ways to achieve that goal. Eyeballing every line of code has never been the most effective way to validate a chance to a piece of software. Recent articles - Conceptual integrity and counting lines of code - 19th August 2026 - Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things - 16th August 2026 - Now we have a timeline of the OpenAI accidental attack against Hugging Face - 7th August 2026

Items scoring 7/10 or above from 11 sources, scored by claude-haiku-4-5-20251001 on relevance to my interests. At most 3 per source.

Scoring categories & sources
  1. Local Wexford or South East Ireland news
  2. Irish or EU-wide affairs affecting citizens broadly: elections, new laws or policy being debated, cost of living, education — especially impacts on mid-life adults or teenagers. Never courts/crime stories.
  3. Irish news on a topic relevant to my interests
  4. Work and tech topics: networking, AI, Kubernetes, platforms, SaaS
  5. AI news including critical or anti-AI perspectives
  6. Gaming: PC gaming, indie gaming, retro gaming
  7. General interests: gardening, woodwork, cycling, fitness, travel
  8. Comics

Sources: Breaking News Ireland, Wexford Local, Hacker News, r/gaming, r/pcgaming, r/antiAI, r/indiegaming, Lenny's Newsletter, One Useful Thing, Newcomer, Simon Willison