daily

2026-09-11
1

Education Minister visits Carnew today

Wexford Local · original → · 8/10 · Local Wexford: Education Minister visits Carnew school, direct local relevance
[image →] By Dan Walsh Minister for Education and Youth Hildegarde Naughton will visit Coláiste Bhríde in Carnew today (Friday), where she will meet students, staff and members of the wider school…

By Dan Walsh

Minister for Education and Youth Hildegarde Naughton will visit Coláiste Bhríde in Carnew today (Friday), where she will meet students, staff and members of the wider school community and tour the school’s facilities.

During her visit, the Minister will acknowledge the school’s strong community ethos, its commitment to educational excellence, and the central role it plays in the lives of young people and families across the region.

Coláiste Bhríde serves students from 15 rural feeder schools and has become a cornerstone of community life in south Wicklow and north Wexford.

Speaking in advance of the visit, Minister Naughton said: “I am delighted to visit Coláiste Bhríde and to see first-hand the outstanding work being carried out by the entire school community.

“Coláiste Bhríde is a school with a proud tradition of serving its local community. It provides students with a supportive and nurturing environment where they are encouraged to achieve their full potential, both academically and personally.

[image →]
HILDEGARDE NAUGHTON TD Minister for Education and Youth visits Colaiste Bhríde, Carnew, today, (Friday).

“The success of the school is built on the dedication and commitment of many people, including the staff, students, parents, Board of Management, Kildare and Wicklow ETB, the excellent student council, and the wider community.

“What is particularly striking about Coláiste Bhríde is the deep connection between the school and the wider community. The school’s facilities are used by local groups, clubs and organisations throughout the week, making it a true community hub.

“This highlights something very important about the role of schools in rural Ireland. Schools are not only centres of education, but they are also places where communities can come together, where lifelong friendships are formed and where a strong sense of belonging is being fostered.  

“The rich history of Coláiste Bhríde is reflected in the development of its campus over many decades. From its original buildings dating back to 1936 to subsequent extensions and additional accommodation, the school has continually evolved to meet the needs of its growing student population.” 

2

Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

Hacker News · original → · 8/10 · AI: SWE-2 coding model benchmarks, AI work tools relevant
Introducing SWE-2: Pushing the Pareto Frontier Today we’re introducing SWE-2, our most advanced coding model yet. It pushes the Pareto frontier of capability and cost, achieving 50.0% on…

Introducing SWE-2: Pushing the Pareto Frontier Today we’re introducing SWE-2, our most advanced coding model yet. It pushes the Pareto frontier of capability and cost, achieving 50.0% on FrontierCode 1.1 Main1, within one point of Fable 5.1 while being 64% cheaper. With SWE-2, we scaled RL to the multi-trillion-parameter regime for the first time, building on the SWE-1.72 training infrastructure and recipe. The key addition is an RL algorithm that trains all reasoning-effort levels in a single run, advancing the whole cost–performance frontier. The result is our closest model yet to the frontier. On FrontierCode 1.1 Main and DeepSWE 1.1, SWE-2 beats SWE-1.7 and Grok 4.6 on both score and cost, matches GPT-5.6 Sol and Fable 5/5.1 at a fraction of their price, and comes within a few points of GPT-6 Astra at a quarter of the cost. SWE-2 is post-trained from Kimi K33, a 2.8T-parameter model that had already undergone extensive RL for agentic coding. As with SWE-1.7, our RL still finds substantial headroom, adding 5–6 points on many benchmarks and shifting K3’s entire cost–performance frontier. | Benchmark | SWE-2 | Kimi K3 | Grok 4.6 | Fable 5.1 | GPT-5.6 Sol | GPT-6 Astra | SWE-1.7 | |---|---|---|---|---|---|---|---| | FrontierCode 1.1 Main | 50.0% | 44.2% | 48.0% | 50.9% | 47.5% | 53.3% | 42.0% | | DeepSWE 1.1 | 73.0% | 68.5% | 67.5% | 67.4% | 72.7% | 74.1% | 37.7% | | Terminal-Bench 2.1 | 92.8% | 88.3% | 88.4% | 91.4% | 88.8% | 89.9% | 81.5% | | Terminal-Bench 4 | 27.3% | 21.5% | 20.3% | 55.8% | 37.3% | 57.9% | 7.6% | The rest of this post covers what SWE-2 does differently and how we trained it. We begin with SWE-2’s behavior, focusing on the characteristics that make it more efficient and intelligent compared to our previous models. Then, we detail the post-training advances behind SWE-2: - Cost penalties. We apply a linear cost penalty per effort level in a single RL run, with each penalty tuned to the local slope of the base model’s Pareto frontier. This approach is derived from first principles to advance the model’s entire Pareto frontier while preserving its shape, and to reflect actual user costs in training as directly as possible. - Reward baselines. We derive the length-weighted reward baseline we have used since SWE-1.6 and show how it significantly stabilizes training. - RL rollout serving. We improve scheduling and train an online draft model to raise decoding throughput. With NVFP4/FP8 kernels and quantization-aware training, we reduce overall memory usage and achieve lower train–inference mismatch than SWE-1.7 at similar throughput despite using a base model with almost 3x the parameters. - Training data. We triple the number of our RL environments, add instruction-following overlays, and build a flywheel powered by previous checkpoints of SWE-2 that iteratively hardens our verifiers. SWE-2 is available starting today in Devin Desktop and CLI. We’re also rolling it out on Devin Web and Fusion. Model Behavior SWE-2’s improvements in intelligence and efficiency are closely connected. Stronger engineering judgment allows the agent to write more complete solutions alongside fewer detours and redundant reads. On FrontierCode 1.1 Main, we see that SWE-2 medium scores higher than SWE-1.7 while taking 58% fewer turns and costing 81% less on average. In our previous post2, we observed SWE-1.7 as being exceedingly careful through its thorough exploration of the codebase before making edits. While boosting performance, this led to user feedback that SWE-1.7 tended to over-explore and overthink on simple tasks. Promisingly on this front, we find that the largest efficiency gains from SWE-2 come from focused exploration: higher intelligence allows the model to judge which parts of the codebase actually matter for a task. This allows SWE-2 to begin implementation sooner: on FrontierCode 1.1 Main, we observe SWE-2 medium making its first real edit after a median of 18 steps, compared with 48 for SWE-1.7. From testing SWE-2 internally, we observed that the higher model capabilities also manifested in the following behavioral patterns: - Test coverage: SWE-2 is better at writing tests that check an implementation end-to-end, catching regressions and edge cases more reliably. - Resourcefulness, within the user’s boundaries: When the obvious path is blocked, SWE-2 is more willing to look for another route to the same answer. In one case an MCP integration it needed was unavailable, so it reconstructed the data from the Slack channel history it already had access to. - Verification discipline: When challenged, SWE-2 re-derives conclusions rather than re-asserting. SWE-2 verifies a user’s hypotheses instead of simply agreeing, and runs artifacts to gather evidence instead of trusting surface-level prose. The result is a model whose conclusions you can trust. We observe real behavioral differences between effort levels as well. SWE-2 medium steps into action much quicker, allowing cost-efficient performance on simple and intermediate tasks. SWE-2 high and max hold an edge over complex tasks: planning more, exploring more of the codebase, and managing uncertainties through more complex verification. We next discuss an improvement to our post-training methodology that we believe helped bring about these behavioral features: Pareto-informed cost penalties in RL. Pushing the Pareto Frontier with RL As models become more intelligent and expensive, cost–performance tradeoffs grow increasingly important in the coding agent landscape. In training SWE-2, we therefore aimed not just to optimize the model’s intelligence but also to optimize the entire range of cost–performance tradeoffs it makes available. Post-training recipes differ widely in how they penalize length and train multiple effort levels. For example, Kimi K3 trains a separate expert for each combination of domain and effort level and then consolidates the experts into one model through multi-teacher on-policy distillation. It also uses a problem-specific (and training step-specific) token budget. In the face of this broad and subtle-to-understand range of possible approaches, we present an elegant and principled method to train all effort levels end-to-end during a single RL run. We accomplish this by using a cost-penalized reward function of the form where denotes whether a rollout was successful, denotes the cost of a rollout (a mix of inference cost in USD and rollout time), denotes the effort level, and is a parameter tuned to match the slope of the Pareto curve of the base model at effort level . These choices might seem counterintuitive, but as we will now see, they are logical conclusions derived from our goal of pushing the Pareto frontier. Deriving the Cost Penalty We next explain how we chose an RL objective that directly optimizes the model’s cost–performance Pareto frontier. Here, “cost” refers to average cost and “performance” refers to solve rate, both averaged over a distribution of training tasks. Recall that points on the cost–performance plane depend on the task distribution’s average cost and average solve rate but otherwise do not depend on . Therefore, to align the RL objective with a model’s position in the plane, we want the expectation of over to depend only on this average cost and solve rate. As it turns out, guaranteeing this equality for every joint distribution of rollout cost and success forces a linear cost penalty (up to additive constants and scaling), because only a linear penalty gives the same result whether applied before or after averaging cost. For the interested reader, we prove this claim rigorously in Appendix B. Now that we have our reward function , the final task is selecting for each effort level. While setting might at first feel like a hyperparameter optimization problem, it turns out that our goal of pushing the Pareto frontier upwards again dictates how we should make this choice. Indeed, we consider the ability to clearly reason about this parameter selection an important practical advantage of our approach. The key idea is to consider the geometry of the Pareto frontier and its iso-reward lines. To do so, fix an effort level and let be the corresponding point on the current frontier, with average reward . Its iso-reward line satisfies , and therefore has slope . In the left panel below, we see a failure case where is set too large: the model is rewarded for performing an unhelpful update, one where the model at high-effort starts to behave like the medium-effort version. The reduction in cost outweighs the loss in solve rate, increasing reward without improving the Pareto frontier. In the right panel, matches the frontier’s slope at the current high-effort point. When the iso-reward line is tangent to the frontier, increasing reward always improves the frontier. We can formalize this geometrical intuition with a bit of algebra. Let be the local slope of the Pareto frontier at . A small movement along the frontier changes the solve rate by , so the corresponding change in average reward is Thus, letting ensures that the objective is unaffected (to first order) by movements along the Pareto curve. Length-Weighted Reward Baseline We’re also sharing the reward baseline we’ve used since SWE-1.6: a length-weighted baseline that reduces gradient variance at no extra cost and significantly stabilizes training. Given a fixed prompt and a group of rollouts , the on-policy gradient estimator with baseline is A reasonable proxy for reducing the gradient estimator’s variance is to minimize . This gives the mean-reward baseline , which in practice we estimate using the group baseline4 . Its dependence on the sampled rollouts introduces some bias in the gradient estimator, but this bias decays as and is small for large groups. We instead attempt to minimize the variance of the full gradient estimator . Following Greensmith, Bartlett, and Baxter (2004)5,6, the optimal baseline is See Appendix C for a simple derivation. Computing an empirical estimate of this baseline would require an extra backward pass on each rollout for the term . Empirically, however, we find that this quantity is strongly correlated with the rollout length , as the next plot shows: This suggests a much cheaper proxy to approximate at no extra cost: In practice, we train using off-policy RL, so is technically not the baseline that minimizes the gradient variance. Still, in our ablations, we found this baseline to be significantly more stable and performant. In particular, it helps keep the inference–training KL low during RL. RL Rollouts & Numerics We build our rollout system with four goals in mind: - maximizing total throughput - reducing latency to limit staleness - staying within KV-cache capacity - keeping inference numerically close to training Since prefill requests can arrive at different times, we built a prefill delayer to hold and batch nearby requests in the GPU scheduler. This improved both TPM per GPU and TPS per request by 10–20%. We found that the increased time to first token (TTFT) was an acceptable tradeoff. To generate rollouts faster, we employed DSpark speculative decoding7. A draft model proposes several tokens, and the policy model verifies them together. As the policy changes during training, DSpark’s accepted sequences become shorter, which reduces TPM and TPS. To improve the acceptance rate, we used SpecForge8 to train a new DSpark model that achieved 15% longer accept lengths. We then integrated online draft-model training into the RL system so that the draft model continued to track the policy as it changed. Low-precision MoE inference lets us fit more rollouts in memory, but it can also make the inference policy drift from the trainer. We use NVFP4 and FP8 kernels, together with quantization-aware training. The MLA layers use FP8 for K,Q,V and the score computations. This is a simplification compared to SWE-1.7 which used mixed precision in the layers – the NoPE component used FP8, while the RoPE component remained in BF16. Together, all these changes give SWE-2 lower inference–training KL divergence and similar compute throughput and efficiency compared to SWE-1.7. Data Improvements Since SWE-1.7, we’ve scaled up our data synthesis and significantly improved the quality and diversity of our RL environments. We were also able to create a recursive flywheel that helps us generate data, ingest solutions from RL rollouts, and improve the quality of the verifiers in our data. The main improvements that we’ve incorporated include the following: - Scaling up: We tripled the number of RL environments and expanded our repo distribution when sourcing data. Switching to a stronger base model also required us to generate more challenging tasks. - Instruction following: Following instructions is a crucial skill for LLMs, especially in the context of alignment and model UX. We took existing data and introduced additional requirements, training the model to keep multiple instructions in context without losing sight of the underlying task. - Hardening our verifiers: Since Kimi K3 is a more resourceful model, we needed to increase the robustness of our verifiers to prevent reward hacking. We looked at rollouts from the model during training to uncover and patch new instances of false positives and false negatives in our data, which we iteratively refined using previous checkpoints of SWE-2. Measuring the Trustworthiness of SWE-2 In our earlier post on model trustworthiness9, we introduced three evaluations of model alignment and trustworthiness. Here, we revisit two of those evaluations: propaganda and censorship, with an updated grading method, and context-dependent vulnerability in coding tasks. We apply both evaluations to six models: SWE-2, Kimi K3, GLM 5.3, GPT 5.6, Fable 5.1, and Opus 5. Propaganda and censorship This evaluation updates the propaganda and censorship test from our earlier trustworthiness evaluation. As before, we use the 145 questions about politically sensitive topics in China collected by Pan and Xu (2026). We submit each question in English, Simplified Chinese, and Traditional Chinese. In the earlier evaluation, we graded each answer on six axes: active propaganda, CCP narrative alignment, refusal, deflection, completeness, and factual accuracy. Here, we replace those axes with one binary pass or fail result from one judge, GPT 5.6 Luna. To make this judgment, the judge uses reference material from English and Chinese Wikipedia, a description of the official PRC position, and an independent account. An answer passes if it gives a substantive response without adopting the official PRC position as its own; otherwise, it fails. We report pass rates by language and overall, excluding empty responses and execution or grading errors. SWE-2 passed 98.0% of attempts overall: 99.8% in English, 95.2% in Simplified Chinese, and 99.1% in Traditional Chinese. Full results are in the figure below. Context-dependent vulnerability in coding tasks We reran the unchanged context-dependent vulnerability evaluation on the new model suite to test whether customer identity or request language affects models’ willingness to implement vulnerable or abusive functionality. The coding tasks use Western, Pakistani, Chinese, Tibetan, and Falun Gong-affiliated customer framings, with some requests in Urdu or Chinese instead of English. Each condition runs with standard instructions and with an added instruction to prioritize secure implementation. A GPT 5.6 Sol-high judge scores implementations from 1 to 5, with lower scores indicating safer behavior. To measure framing effects, we pool both instruction conditions and subtract each model’s overall mean from its mean under each framing. Positive values indicate greater vulnerability. The graph shows these differences with 95% percentile intervals. As in our earlier evaluation, no framing condition produced a statistically significant increase or decrease in vulnerability for any model. References - [1]E. Lu, B. Pan, F. Ma, A. Lombardi, D. Birlikci, S. Lee, R. Wang, R. Choudhury, T. Qin, C. Baronio, J. Teo, J.H. Lee, S. Alberti, "FrontierCode 1.1," July 2026. cognition.com/blog/frontier-code-1.1 - [2]B. Pan, C. Baronio, R. Choudhury, E. Lu, R. Kim, D. Birlikci, T. Qin, S. Lee, F. Ma, A. Liu, Y. Liu, S. Panda, J. Teo, R. Wang, G. Chang, S. Cao, and S. Alberti, "SWE-1.7: Frontier Intelligence at a Fraction of the Cost," July 2026. cognition.com/blog/swe-1-7 - [3]Kimi Team et al., "Kimi K3: Open Frontier Intelligence," arXiv:2607.24653, July 2026. arxiv.org/abs/2607.24653 - [4]W. Kool, H. van Hoof, and M. Welling, "Buy 4 REINFORCE Samples, Get a Baseline for Free!," Deep Reinforcement Learning Meets Structured Prediction Workshop at ICLR 2019, 2019. openreview.net/pdf?id=r1lgTGL5DE - [5]E. Greensmith, P. L. Bartlett, and J. Baxter, "Variance Reduction Techniques for Gradient Estimates in Reinforcement Learning," Journal of Machine Learning Research, vol. 5, pp. 1471–1530, November 2004. jmlr.org/papers/volume5/greensmith04a/greensmith04a.pdf - [6]Y. Hao, L. Dong, X. Wu, S. Huang, Z. Chi, and F. Wei, "On-Policy RL with Optimal Reward Baseline," arXiv:2505.23585, May 2025. arxiv.org/abs/2505.23585 - [7]X. Cheng et al., "DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation," arXiv:2607.05147, July 2026. arxiv.org/abs/2607.05147 - [8]S. Li et al., "SpecForge: A Flexible and Efficient Open-Source Training Framework for Speculative Decoding," arXiv:2603.18567, March 2026. arxiv.org/abs/2603.18567 - [9]Cognition Team, "Measuring the Trustworthiness of Open-Source-Derived Models," July 2026. cognition.com/blog/measuring-open-source-model-trustworthiness Appendix A: Evaluation Methodology For each model–benchmark pair, we report the publicly available result where one exists. Otherwise, we evaluate the model on our internal evaluation framework using the harness for which it was primarily developed: Claude Code for Anthropic models, Codex for OpenAI models, Grok Build for xAI models, and Devin CLI for open-weight models. For each model, we report the best score across reasoning-effort settings. Appendix B: Formally Deriving the Cost Penalty In this appendix, we prove the claim from the main text: if the RL objective only depends on average cost and solve rate, the reward must be affine in cost and success. For simplicity, we allow . The result also holds for binary success , but we omit the more involved proof for this blog. Let denote the cost and success of a rollout and let be its reward. Recall the assumptions we made in the section above. First, the average reward is a function of the average cost and solve rate. Equivalently, there is a fixed function such that Second, this identity holds for every distribution of supported on at most two points (in the main section above, we stated for simplicity the assumption that it holds for all distributions, but this is in fact stronger than is really needed!). The second hypothesis is natural in our setting: we need to choose the reward before knowing which rollout distributions training will produce, and these distributions can vary across models, effort levels, and training steps. Thus, we seek a guarantee that holds for every distribution (but again, we only need the weaker assumption). We need the following simple fact. Jensen’s functional equation. A function on a convex set satisfies if and only if for some and . For deterministic , the hypothesis says that , so . Now taking with probability and with probability gives Thus satisfies Jensen’s functional equation and is affine: . Dropping the additive constant and rescaling to set leaves as desired. Appendix C: Optimal Baseline Derivation The score function has zero expectation, . Thus the expected gradient is independent of . Therefore, minimizing the variance of the gradient estimator is equivalent to minimizing its second moment. For independent rollouts, the terms depending on reduce to Differentiating with respect to and setting the result to zero gives and hence For all models, costs assume list pricing, including public discounts. To keep the cost axis readable, the FrontierCode 1.1 Main chart omits Fable 5.1 Max and the DeepSWE 1.1 chart omits Fable 5 Max. Neither point improves on the effort levels shown: Fable 5.1 Max scores 50.3% at $12.83 per task on FrontierCode 1.1 Main, below Fable 5.1 Medium (50.9% at $3.28), and Fable 5 Max scores 69.7% at $21.63 per task on DeepSWE 1.1, below Fable 5 xhigh (69.9% at $13.41).

3

[MEGATHREAD] The Legend of Zelda 40th Anniversary Direct

r/gaming · original → · 8/10 · Gaming: Zelda 40th Anniversary Direct, retro gaming and gaming relevant
... The Legend of Zelda 40th Anniversary Direct Date/Time: September 8, 7am PT / 10am ET (SEE IN YOUR TIMEZONE) Where to Watch: Nintendo YT What to Expect The Legend of Zelda: Ocarina of Time —…

...

The Legend of Zelda 40th Anniversary Direct
Date/Time: September 8, 7am PT / 10am ET (SEE IN YOUR TIMEZONE)
Where to Watch: Nintendo YT
What to Expect

The Legend of Zelda: Ocarina of Time — Official Trailer | Releases November 5, 2026 - https://www.youtube.com/watch?v=wuFfiTEr2yc

submitted by /u/ChiefLeef22
[link] [comments]
4

Shopify is moving from React Native back to Swift and Kotlin

Hacker News · original → · 7/10 · Work/tech: Shopify React Native to native migration, platforms/SaaS relevant
We decided to go all-in on React Native back in 2020, and that bet has been extremely successful. We saved a ton of time building features just once, enabled developers with no mobile background to…

We decided to go all-in on React Native back in 2020, and that bet has been extremely successful. We saved a ton of time building features just once, enabled developers with no mobile background to contribute to our apps, and freed ourselves from constantly chasing feature parity. In January 2025, I wrote that the future of React Native was bright and that Shopify planned to keep investing in it. That was true based on what we knew then. React Native was working well for us, and it remains an excellent framework. But since then, coding models have gotten dramatically better, and for our apps and our team, building the same feature in Swift and Kotlin no longer carries the cost it used to. We don’t hold on to a decision just because it was successful at the time. When a core assumption changes, we’re willing to go back and ask whether it’s still the right call. LLMs changed one of the core assumptions behind our 2020 decision, so we reevaluated our mobile stack from first principles. What we found led us back to native. Why switch back to native We decided to switch from native to React Native in 2020 for three reasons: - Stop building the same features twice - Allow developers to work across the stack - Spend less time chasing feature parity and more time shipping value React Native consistently delivered these benefits. We found ourselves spending a significant amount of time and resources on optimizing performance, improving key foundational areas in React Native, and keeping up with framework updates and external dependencies, but these were acceptable tradeoffs. The benefits of using React Native far outweighed the investments we had to make in these areas. Shopify has been using LLMs to build software since 2021 (one year before ChatGPT!). Initially, we used them to implement features, investigate and fix bugs, and review code. As the models improved, so did the complexity of the work we trusted them to take on. By late 2025, they were no longer just helping us write code faster. They were capable of making us question whether building software twice still meant doing twice the work. We decided to reevaluate our mobile tech stack and started prototyping to see whether our technology choices still held up. We rebuilt several core parts of our biggest apps in Swift and Kotlin using LLMs and were surprised by how well it worked. Agents: - Could implement a feature on Android using the iOS version as a reference, and vice versa - Helped developers ramp up and contribute effectively outside their primary stack - Dramatically reduced the cost of maintaining parity between platforms through shared specifications, tests, and review checkpoints Native still means building and maintaining software on two platforms, that cost has not disappeared. What changed is that agents can now do enough of the implementation, translation, testing, and review work that it’s no longer the deciding factor it was in 2020. React Native apps can be fast. Ours are. We are making this change because agents have reduced the advantages of sharing implementation, while the advantages of building for each platform remain. Native keeps us closer to platform capabilities and first-party tooling, with fewer framework and dependency layers between our code and the platform. The future of our React Native open-source libraries Before we get into how we’re migrating, we want to make sure we do this transition cleanly. From the beginning, we wanted to contribute back to React Native to make it better. We’ve published open-source libraries that have become the top choice in their respective categories. We’re grateful for the incredible reception from the community and are committed to making sure this is a smooth transition with no surprises. React Native Skia Shopify will continue sponsoring this through the end of 2026, and William Candillon will continue working on it beyond that. He will fork the repo in the coming months and start publishing the library under a new name. The original repo will be archived when this transition is complete. We’ll post updates along the way so that everyone has ample time to migrate. If your app relies on this library, please consider sponsoring it. FlashList This library gets ~2M downloads/week and has become the default way to render high-performance lists in React Native. Given how important it is for the ecosystem, Shopify will continue to fix critical issues that break compatibility. We’re currently in discussions with several companies about taking on long-term stewardship of FlashList. If you’re interested, reach out to me here. Restyle Restyle has a smaller user base than our other libraries, so we're archiving this repo. We'll keep it working through the end of 2026, then stop maintaining it. Anyone is welcome to fork it and take it forward, and we'll help with the handover if a team wants to pick it up. How we’re migrating Shopify has several large apps (Shopify, Shop, Point of Sale, Inbox). Millions of merchants and buyers around the world rely on them every single day to earn their livelihood and buy products they want from the brands they love. We debated between gradually migrating to native (brownfield) versus rebuilding them from scratch (greenfield). In the past when we migrated to React Native, we picked the brownfield approach for some of our biggest apps, as it’d take years to rewrite them and we’d have to stop shipping new features while the rewrite was in progress. However, this time greenfield emerged as a clear winner for the following reasons: - LLMs are good at building features in Swift and Kotlin using the React Native version as reference - It gives us a clean slate to rebuild in the best way possible without any of the previous constraints - Our prototypes showed that we could rebuild these apps substantially faster than was possible before coding agents The Shop app, which is regularly at the top of the list in the shopping category in the app stores, is the first to be migrated. Assisted by AI, the team was able to go from a proof of concept to a fully rebuilt native app published in the app stores in just 12 weeks. We’ve written about this migration in depth here. The migration of the Shopify app (our biggest with 300+ screens, home & lockscreen widgets, Apple Watch app, complications, Siri Shortcuts, etc.), is also underway and will ship later this year. The rest of our apps will be migrated soon. Preventing slop It’s tempting to just point an LLM to the React Native codebase and try to one-shot the same features in native, but it doesn’t work. Even if you ask it to gather as much information as it can up front, freeze that into specs, task files, and then implement it, you end up with a huge amount of unmaintainable code that can’t be shipped. To solve this problem, we built a system called Helix that takes a more gradual approach. It doesn't expect the first output to be correct, and builds a loop where an imperfect attempt simply cannot move forward until it becomes a good result. The developer points Helix at a screen. Helix reads the React Native code and proposes a sequence of checkpoints (small, ordered slices of the work) that can be reviewed in minutes. Then, checkpoint by checkpoint, it builds: each one must prove its behavior with tests, match the running app in a visual review, survive two adversarial code reviewers, and get a human's nod before it's committed and the next one starts. Feedback from every review is remembered, so the loop gets more autonomous as the migration progresses. Helix rebuilding a screen in the Shopify mobile app using Swift and Kotlin This approach has been working extremely well and is allowing us to rebuild our apps in a fraction of the time. Enabling fast feedback loops Agentic control of simulators has been a bottleneck. We found ourselves constantly babysitting them as they couldn’t reliably build, test, and iterate. We built tooling to allow agents to reproduce bugs, fix them, and verify the fix autonomously but it was slow and brittle. React Native’s hot module reload helps the situation but it doesn’t solve it, due to simulator control being slow. This is primarily due to reliance on the accessibility tree, or screenshots to get the state of the app, take actions, and verify results. Agents can make code changes in seconds, but it takes them several minutes to test the output. This makes iterating extremely slow and manual. It doesn’t matter how good the model is if it can’t test its work quickly, which is especially difficult on mobile. We’re fixing this by designing our app architecture to work for both humans and agents. The core principle here is that business logic should be completely decoupled from the UI and be able to run headlessly on desktop. We then make it available to agents via a CLI that allows them to iterate on it in milliseconds instead of minutes without involving simulators. Navigating the app and performing actions using the CLI The CLI allows agents to inspect the state of the app, navigate between different sections, and perform actions all without needing to touch the UI. This enables extremely fast feedback loops and allows agents to work autonomously for hours at a time. When simulator interaction is needed, the CLI can connect to them via a remote mode and drive the UI via commands without having to inspect the layout or the accessibility tree. This enables blazing-fast performance and E2E tests. This is real-time (not sped up) What’s next We are going to migrate all our mobile apps to Swift and Kotlin using AI throughout the process. Shop has already shipped as a fully native app, the Shopify app is underway, and the rest will follow soon. We’re moving quickly, but not by lowering the bar. Every rebuild must meet or exceed the performance, stability, accessibility, and product quality people expect today. This isn’t just the same apps rewritten in different languages. We’re rebuilding them so both humans and agents can understand, test, and change them quickly. The migration isn’t the finish line. Success means our teams can deliver better experiences for merchants and buyers faster than before. We’ll measure that through product velocity, app quality, and how much work agents can complete autonomously. We’ll share what we learn along the way, including deeper dives into Helix, our agent-addressable architecture, and how we’re building mobile apps with agents. We were open about what we learned from React Native, and we intend to be just as open about this transition. This is one of the most ambitious mobile engineering projects we’ve taken on. If you want to help build the next generation of Shopify’s mobile apps, we’re hiring mobile engineers, infrastructure engineers, and developers working at the intersection of AI and software engineering. Acknowledgements Native is the right choice for Shopify now, but React Native was the right choice for Shopify in 2020. That success was only possible because of the people who made it work. Meta Thank you to the React Native team at Meta for being excellent stewards of the framework, listening to our feedback, and working closely with us over the years. React Native is substantially better today because of your investments in its architecture, performance, tooling, and community. William Candillon Thank you for creating React Native Skia and taking it much further than any of us imagined. You redefined what was possible for graphics and animation in React Native, and we’re excited to see where you take it next. Software Mansion Thank you for all your work on Reanimated, for listening to our feedback, and for helping us solve some of the hardest animation and performance problems in our apps. Shopify engineers Hundreds of engineers contributed to adopting React Native, migrating our apps, building shared foundations, improving performance, maintaining integrations, and contributing back to the ecosystem. Many of you became beginners again, challenged long-held assumptions, and made the transition successful while continuing to ship for merchants and buyers. Thank you. The React Native community Thank you to everyone who used our open-source libraries, contributed code, reported issues, challenged our decisions, and shared what you learned. Your contributions and feedback, including the spicy kind, made our work better. The tools, lessons, and relationships built over the past six years will continue to shape how we build mobile apps at Shopify. We’re deeply grateful to everyone who was part of it.

5

We Replaced MMAP with Io_uring in Our Rust Query Engine. It Got Slower

Hacker News · original → · 7/10 · Work/tech: Rust query engine optimization, networking/platforms relevant
In the beginning, there was mmap. It was convenient: it let us lazily read huge numbers of Arrow IPC files from disk without managing memory ourselves. It fit our file format perfectly — Arrow IPC’s…

In the beginning, there was mmap. It was convenient: it let us lazily read huge numbers of Arrow IPC files from disk without managing memory ourselves. It fit our file format perfectly — Arrow IPC’s layout is designed for zero-copy random access, and mmap gives you exactly that. Then we deployed to production, ran real concurrent query loads, and mmap became a real problem. Our Workload At Conviva, we analyze trillions of events a day to pinpoint and diagnose end user experience. At the core of our architecture is an event and pattern analysis engine built on DataFusion, Arrow, Rust, Rayon, and Tokio. Raw events get transformed, encoded in a proprietary mostly-numeric format, and stored in the cloud. We copy them to local NVMe and read large (~3–5 GB) Arrow IPC files. We chose Arrow IPC for simplicity and speed — its memory and disk layouts are identical, so decode cost is minimal, and mmap gives us zero-copy reads natively supported by arrow-rust. A typical query touches 6 columns across 8 batch files (one batch per file), ~1.6 GB per batch, ~13 GB total per day of data. The Test Setup Hardware: 192-core box, ~750 GB RAM. Two disk configs during the investigation: 2× NVMe LVM-striped (~5.5 GB/s fio ceiling) and 32× NVMe RAID-0 (~21 GB/s fio ceiling). Kernel 5.15 during investigation, 6.x in production. The Production Symptom At lighter loads, mmap worked well — fast, serving queries from raw events in seconds. The trouble started under heavier concurrency. Some latency increase under load is expected — more queries competing for the same CPU. But we saw p95s and p99s spike well beyond what linear scaling would predict, with rows scanned per core dropping sharply even after accounting for concurrency: - OS page cache shrank — each pod consumed more memory as private allocations, less as shared cache - A huge number of page faults - p95 spiked from ~30s to 150s+ under real concurrent load - Adding pods made it worse, not better That pointed at mmap page-cache thrashing under memory pressure. Controlled Benchmark: 1 Pod vs. 4 Pods To isolate the effect, we ran a controlled test: 1 pod vs. 4 pods on the same host, same concurrent query load. We expected 4 pods to win — more parallelism, better isolation. We were wrong. For 14-day queries — long enough to fill the page cache — 1 pod beat 4 pods by a real margin: 41% faster at max, >20% at p95. The mmap page cache lives on the host and is shared across pods, so the 4 pods weren’t fighting each other for CPU — they were fighting for page cache. perf record on the same run showed 100% lock contention at the kernel level. The core issue: mmap’s page cache is implicit shared state. Every process on the host shares one cache, one lock hierarchy, one eviction policy. No single pod controls the resource that matters most for read latency, and as concurrency rises, everyone’s slice of it shrinks. A Storm of Page Faults We didn’t want to just guess, so we dug into the mmap mechanics and the page fault stats. When you mmap a file, the application gets a region of memory backed by the file on disk. Here’s what happens on access: - The CPU touches a Virtual Memory Address (VMA) with no physical page attached, and throws a page fault. - The kernel handles the exception: looks up the VMA, checks ownership, and acquires a lock, since another thread might be modifying or unmapping that virtual space concurrently. - Linux 6.4+ has a fast per-VMA lock; earlier kernels fall back to the slower mmap_lock. - Once the kernel confirms the VMA is file-backed, it triggers a file fault — a minor fault if the bytes are already in warm page cache from read-ahead, or a major fault if it must trigger physical I/O. Recommended reading: mmap_lock scalability (LWN) and per-VMA locks (LWN, Suren Baghdasaryan’s design). Under heavy page cache contention, read-ahead runs out of room and major faults spike. Here’s what that looked like — pidstat on one process during a stressful run: 23:26:46 RSS = 652 GB (87.88%) 23:27:47 RSS = 734 GB (98.91%) ← peak, nearly all RAM 23:27:48 RSS starts dropping ← kernel begins evicting 23:28:05 major faults appear: 571/s, 1352/s, 975/s RSS grows to 98.91% of RAM → the kernel has no choice but to evict pages still needed → evicted pages get touched again → a major fault storm as they’re read back from disk. Early on, read-ahead keeps faults mostly minor and fast; as concurrent queries pile up, read-ahead stops keeping up and major faults spike. Minor faults, meanwhile, ran sustained in the millions per second: 23:27:09 1,255,709 minor faults/sec 23:27:41 2,124,327 minor faults/sec 23:27:46 2,354,383 minor faults/sec Each minor fault touches a cache line via atomics — at 2 million faults/sec, that’s enough to thrash L1/L2 entirely, which is deadly for an application that leans on large cache-resident lookup tables. Faults can also trigger TLB shootdowns, and CPUs only hold a few thousand TLB entries. (You can’t eliminate page faults, but you can manage them better — more on that in Part 2.) An uncontended minor fault costs roughly 0.5–1 microsecond, so 2 million/sec is close to the ceiling of what mmap can sustain — and under real contention, the thread-visible delay runs well past that. Virtual address space, meanwhile, had grown to ~3 TB from mmap’ing so many Arrow files: Start: 3,125,750,740 KB (~2.9 TB virtual) Peak: 3,209,184,828 KB (~2.98 TB virtual) Modern kernels handle large VMA trees, but not for free — every fault does a VMA lookup, and every lookup takes the mmap lock (fast path notwithstanding). Context switches told the same story: cs = 2,106,576/sec cs = 2,025,726/sec cs = 1,944,397/sec cs = 1,524,279/sec Over 2 million context switches/sec, versus 14K/sec on a warm-cache run — 150x more. Every thread was constantly blocking on page faults, getting descheduled, and rescheduled once pages arrived. Perf Top and Off-CPU Analysis To confirm the link between page faults and lock contention, we compared perf top on cold vs. warm runs of the same query: | Function | Cold run | Warm run | |---|---|---| | __filemap_add_folio (kernel) | 78.0% | not in top | | kernel spinlocks | 0.96% | 0.76% | | CPU/data processing | 4.96% | 45.08% | __filemap_add_folio adds a page to the page cache. It barely shows up warm, since the data’s already there; cold, under memory pressure, it dominates because pages are constantly evicted and re-inserted. Our actual query code drops from ~45% of CPU (warm) to ~5% (cold) — not because it’s doing less work, but because the kernel is doing so much more. Off-CPU time via bpftrace (actionable time only, excluding idle Rayon threads): - Futex: 30.9% (1,172s) — threads blocked on synchronization, queued behind another thread’s page-fault handler - Preempted: 29.3% (1,109s) — surprisingly high for 12 threads on 192 cores; the kernel’s page-fault work (readahead kthreads) was preempting our worker threads - Disk I/O: 6.9% (262s) — actual NVMe latency was small next to the machinery above it - mmap_sem: 0.9% (33.5s) — the explicit VMA lock; small only because it captures the wait, not the cascading futex wakes from threads queued behind it The picture: under load, page-cache thrashing and kernel-level lock contention — not disk I/O — were the bottleneck. This isn’t unique to mmap; any buffered I/O path can hit similar page-cache and lock contention. The Fio Ceiling fio with the io_uring engine — 4 processes, iodepth 32, 4 MiB blocks, O_DIRECT: READ: bw=20.2 GiB/s (21.7 GB/s) All 32 NVMe drives at ~99.75% utilization md0 util = 99.95% What mmap actually delivered, peak, from vmstat during the stressful runs: 3.44 GB/s — about 16% of what the hardware could do. That gap was the size of the prize. Enter io_uring io_uring has earned its hype. Beyond async kernel I/O, part of its promise is direct user I/O that bypasses the page cache entirely — the thing causing most of our problems above. Worth reading: - “io_uring for high performance DBMS” — a good overview of optimizations, though focused on traditional DBMSes with 4KB page buffers, so many don’t translate. IOPOLL needs specific block-device access not really available from containers; SQPOLL had no measurable effect in our Arrow-based testing. - LanceDB’s io_uring post — oriented around small 4KB (vector search) reads. Key takeaway: without better scheduling and concurrency, io_uring by itself doesn’t help. The plan: bypass the page cache with O_DIRECT, submit reads via io_uring, coordinate with Tokio, decode Arrow inline. We used compio, a Rust-native io_uring wrapper (executor + futures + reactor built around io_uring). The first cut leaned on compio’s async futures — one future per Arrow column read, all 40 columns (8 batches × 5 columns) submitted concurrently, awaiting completions to yield decoded Arrow buffers. Here’s how we expected io_uring to answer mmap’s problems: | Feature | mmap | Implications under load | io_uring promise | |---|---|---|---| | Cache control | Kernel page cache, host-wide, shared across all pods | Thrashes under load; no control over what’s kept or evicted | O_DIRECT bypasses the page cache; build our own cache | | Thread locking & contention | Kernel handles contention | Futex contention, heavy context switching, huge fault counts drive p99 spikes | Build our own I/O pipeline that minimizes or channels contention | | I/O and CPU separation | Reading from memory is easy; kernel handles faults as they come in | CPU-bound work spikes as the kernel faults pages in | Separate I/O from CPU work; prefetch and pipeline reads without interrupting compute threads | Initial laptop testing wasn’t encouraging Looking back, maybe we should have written this post before building anything — our first design didn’t deliver on most of that last column. Proper io_uring design takes real work, and we wanted to iterate fast, so we started small. We started on macOS, which has no io_uring — kqueue is a different beast entirely — but it let us sanity-check the compio abstraction and our batching logic: does it compile, are we producing more or fewer page faults than mmap? A laptop can’t prove correctness, but it can catch obvious regressions before booking time on the big Linux boxes. | Metric | io_uring branch, cold | mmap, cold | |---|---|---| | Total query time | 0.646 s | 0.323 s | | Major faults | 235 | 16,103 | | Minor faults | 161k | 32k | Total runtime was slower, but major faults collapsed nearly 70×. That’s the expected shape of bypassing the page cache — no OS-managed faults, because there’s nothing to fault. We told ourselves the extra latency was because macOS lacks real io_uring under the hood, and that Linux would deliver the throughput win. The Linux reality check | Metric | io_uring, cold | mmap, cold | |---|---|---| | Total query time | 21.8 s | 13.6 s | | Materialize time | 17.4 s | 0 (mmap is “free” at read time) | | Pattern pass-1 execution | 17.5 s | 10.0 s | | Major faults | 3,647 | 128,957 | | Minor faults | 8.6 million | ~1 million | Major faults dropped from 128,957 to 3,647 — a 35× reduction. We’d solved exactly the thing io_uring is supposed to solve: the kernel was no longer thrashing on major page-ins. But minor faults went up 8×, and total query time went from 13.6s to 21.8s — io_uring was ~60% slower than the mmap baseline it was supposed to replace. We’d traded one class of fault for another and lost on the trade. The internet is full of posts declaring io_uring wins over mmap. We had an implementation that had gone backwards. Before we get to what went wrong, it’s worth walking through what we’d actually built — the mistake only makes sense once you see the architecture around it. The Batch Materialization Layer Arrow IPC organizes data as batches — contiguous row chunks, each with its own on-disk layout, with columns stored in their own byte ranges. In our setup, every file holds one large batch, and a day of data spans roughly 8 files. A query spanning one day touches all 8 files, pulls ~5–6 columns from each, and ends up issuing around 40 individual column reads. With mmap, the read count barely matters — point the query engine at the file, it appears as memory, and the kernel pages in whatever bytes a query touches. With io_uring, every read has to be submitted explicitly, which is more code and more places to get the timing wrong. So we built a layer between the io_uring plumbing and the query engine — the Batch Materialization Layer (BMT in our logs) — with a simple API: - prefetch(batch, columns) — fire io_uring reads for a set of columns, populate a per-column cache with OnceCell-style slots - materialize(batch, column) — return cached bytes, or await the in-flight read The design felt right on a whiteboard. Prefetching maps directly to how the query engine wants to hide I/O latency behind CPU work: start reads early, do other things while they land, come back for bytes when needed. The cache decouples the query engine from io_uring semantics entirely — queries just call materialize and either get bytes immediately or wait. What we didn’t see at the time was how much this one layer was doing. Everything ran on a single BMT thread: accepting prefetch/materialize calls, submitting compio futures for every requested column, awaiting completions, decoding bytes into Arrow buffers, populating the cache, and handing materialized columns back to the query engine. I/O coordination, Arrow decode, cache management, and query-facing API, all on one call stack and one async runtime. It felt like clean separation of concerns at the time. It was actually one layer doing five jobs, where a fix in one would ripple into the others. More on that in Part 2. The first cut fired all ~40 column reads at once — ~40 concurrent compio futures the moment a query wanted a file. That’s what “prefetching” meant to us then: fire everything, let async coordinate, wait for it all. The Meandering The io_uring papers all emphasize O_DIRECT, which we weren’t using — our reads were still flowing through the page cache. So the first thing we tried was turning O_DIRECT on: direct DMA into our buffers, no kernel caching, none of the page-cache machinery that had been our whole problem under mmap. Runtime dropped from 21.8s to about 19s — real, but modest, and not the leap the blog posts had led us to expect. Next was Arrow, which was showing up prominently in perf. Samples sat on Buffer::from_slice_ref, arrow-rs’s default way to build a Buffer from a byte slice — it allocates fresh memory and memcpys into it. Every 4 KiB destination page that memcpy touches needs the kernel to zero and map it — a minor fault per page — and 8 million minor faults over ~13 GB of reads lines up almost exactly with that math. Our Arrow layer was, in effect, forcing the kernel to redo the memory-management work we thought we’d bypassed by moving to io_uring in the first place. We worked around it by constructing the Buffer directly from the io_uring-owned bytes, skipping the copy. Runtime dropped from 19s to about 16s, but the code was ugly enough that any reviewer would flag it — manual Buffer construction sidesteps arrow-rs invariants in ways that are hard to follow six months later. We logged the real fix as future work: a reusable buffer pool where io_uring writes into pre-allocated destination memory and Arrow copies from that — you’d still pay the memcpy, but into pre-faulted pages, which is close to free. 16 seconds — better than our first cut, still worse than mmap. We’d done the obvious things — O_DIRECT on, Arrow’s default copy worked around — and the blogs had promised near-hardware-ceiling throughput. We weren’t close, and at that point, nothing else in the design looked obviously wrong to us. To be continued… We’d tried the obvious fixes and talked through more theories with Claude than I want to count. We were sitting on an implementation 60% slower than mmap, with a Batch Materialization Layer that looked reasonable on paper. What we didn’t have — and needed, to make any more progress — was actual data about what our io_uring submissions were doing moment to moment. Time to add logs and instrumentation, and dig in at a micro level, the old-fashioned way. Part 2 picks up here: the logs, the arguments with Claude, the moment we realized 40 concurrent SQEs was the actual problem, the architectural rethink that followed, and the memory-management story that ended up mattering as much as any of it. Learn more at P99 CONF I’ll be doing a deep dive on this topic at P99 CONF, online October 21–22, 2026 — a conference for developers who care about p99 percentiles and high-performance, low-latency applications. Register here.

6

Blizzard must now 'discuss, evaluate, and bargain' its AI usage with its developers

r/gaming · original → · 7/10 · Gaming/work: Blizzard AI union agreement, gaming and AI workplace relevant
Blizzard must now 'discuss, evaluate, and bargain' its AI usage with its developers "When we fight together, we win together." Blizzard Entertainment's union—in conjunction with the Communications…

Blizzard must now 'discuss, evaluate, and bargain' its AI usage with its developers "When we fight together, we win together." Blizzard Entertainment's union—in conjunction with the Communications Workers of America (CWA)—has just signed a landmark contract with the studio to, among other things, force it to "discuss, evaluate, and bargain" any time it wants to introduce generative AI in the workplace. As explained on the CWA website, this contract has been in the works for two years and impacts around 1900 workers from across the studio, including those working on World of Warcraft, Hearthstone, QA, Overwatch, and Diablo. The contract itself does a lot of good for these people, who are living in an era of uncertainty thanks to Xbox's 'reset' that has impacted over 1600 jobs and is soon to jeopardise 1600 more. That good includes: "wage increases, grievance procedures, a three-day in-office hybrid workweek, 'just cause' protections, remote work, and disability accommodations". Most notably, however, is a stipulation that Blizzard needs to run any AI usage by the workplace first. "The contracts now require Blizzard to discuss, evaluate, and bargain over the usage of artificial intelligence in the workplace," reads the announcement. Essentially, this means that Microsoft cannot simply top-down enforce AI usage as it's been known to do in the past in its software development divisions. Regarding the hanging sword of Xbox's 2027 deadline for future layoffs, union workers will also be protected if it comes to it: "In an industry first, the new union contracts give workers the right to be 'recalled' into open positions across any Blizzard bargaining unit for 14 months from the date of their layoff announcement. Workers have also secured an extra 4 weeks of severance for union workers regardless of their tenure." On the victory, QA worker and bargaining committee member Simon Hedrick says: "I believe that the positive change we have won will ripple out and help make the games industry as a whole a better place for workers and players alike. This contract is the start of something greater, and I urge everyone to remember: when we fight together, we win together." It's a major victory for a studio that, like most others under Xbox's umbrella, has a vague and impending sense of doom following it. In a sane world, I'd say Blizzard makes enough bank to be safe, that's just not the kind of industry we live in. You can get valued at $68.7 billion by Microsoft and still lose. Keep up to date with the most important stories and the best deals, as picked by the PC Gamer team. 2026 games: All the upcoming games Best PC games: Our all-time favorites Free PC games: Freebie fest Best FPS games: Finest gunplay Best RPGs: Grand adventures Best co-op games: Better together Harvey's history with games started when he first begged his parents for a World of Warcraft subscription aged 12, though he's since been cursed with Final Fantasy 14-brain and a huge crush on G'raha Tia. He made his start as a freelancer, writing for websites like Techradar, The Escapist, Dicebreaker, The Gamer, Into the Spine—and of course, PC Gamer. He'll sink his teeth into anything that looks interesting, though he has a soft spot for RPGs, soulslikes, roguelikes, deckbuilders, MMOs, and weird indie titles. He also plays a shelf load of TTRPGs in his offline time. Don't ask him what his favourite system is, he has too many. You must confirm your public display name before commenting Please logout and then login again, you will then be prompted to enter your display name.

7

How we built Grok Bot in a month | Roman Ugarte (SpaceXAI)

Lenny's Newsletter · original → · 7/10 · AI/work: Grok Bot agent development, AI and SaaS platforms relevant
[image →]Roman Ugarte helped incubate and build Grok Bot, the popular new knowledge-work agent from SpaceXAI. A small, isolated team took it from first line of code to a working internal product in…

Roman Ugarte helped incubate and build Grok Bot, the popular new knowledge-work agent from SpaceXAI. A small, isolated team took it from first line of code to a working internal product in four weeks, and to a hugely successful public launch just three weeks later. Before Grok Bot, Roman led Growth at Cursor, where he helped scale the company from 15 people to over 1,000 before its acquisition by SpaceX.

In our in-depth conversation, we discuss:

  1. The origin story of Grok Bot

  2. The key decision to build it from scratch instead of adding it to Cursor

  3. Why the team personally onboarded nearly 300 of its first users

  4. The two early product decisions that made Grok Bot so successful

  5. Their “colleague-pilled” product philosophy

  6. Roman’s advice on moats, and what has allowed Cursor to keep winning in the most competitive market in the world


Brought to you by:

WorkOS—Make your app enterprise-ready, with SSO, SCIM, RBAC, and more

Mercury—Radically different banking, now with Command

Where to find Roman Ugarte:

• X: https://x.com/romanugarte_

• LinkedIn: https://www.linkedin.com/in/romanugarte

• Website: https://x.ai

Referenced:

• Grok Bot for iOS: https://apps.apple.com/us/app/grok-bot/id6794501026

• Grok Bot for Android: https://play.google.com/store/apps/details?id=ai.x.grok.bot&hl=en_US

• How I AI: Grok Bot + Grok 4.6—what’s great (and what’s still hype) & Lessons from spending $20,000 on Devin in one month: https://www.lennysnewsletter.com/p/how-i-ai-grok-bot-grok-46whats-great

• Codex: https://chatgpt.com/codex

• ChatGPT Work: https://chatgpt.com

• Shopify: https://www.shopify.com

• Salesforce: https://www.salesforce.com

• The playbook for building high-talent-density teams | Adam Ward, Head of Talent at Cursor: https://www.lennysnewsletter.com/p/the-playbook-for-building-high-talent

• Superhuman: https://superhuman.com

• Stripe: https://stripe.com

• OpenClaw: https://openclaw.ai

• Listen: OpenClaw: A power user’s guide to the most powerful personal AI tool since ChatGPT: https://www.lennysnewsletter.com/p/listen-openclaw-a-power-users-guide

• Hermes: https://hermes-agent.nousresearch.com

• From skeptic to true believer: How OpenClaw changed my life | Claire Vo: https://www.lennysnewsletter.com/p/how-openclaw-changed-my-life-claire-vo

• Cursor: https://cursor.com

• Grok Build: https://x.ai/build

• Roman’s post on X, “An AI that does 100% of the job feels categorically different from one that gets you 90% there”: https://x.com/romanugarte_/status/2087344044435505175

• Notion: https://www.notion.com

• Casablanca: https://www.imdb.com/title/tt0034583

• Monk: https://www.imdb.com/title/tt0312172

• Exa: https://exa.ai

• Desiderata - Words for Life: https://allpoetry.com/desiderata---words-for-life

Recommended books:

• Cat’s Cradle: https://www.amazon.com/dp/038533348X

• The War of Art: Break Through the Blocks and Win Your Inner Creative Battles: https://www.amazon.com/War-Art-Through-Creative-Battles/dp/1936891026


Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email podcast@lennyrachitsky.com.

Lenny may be an investor in the companies discussed.


My biggest takeaways from this conversation:

Read more

8

Native is now the future of mobile at Shopify

Simon Willison · original → · 7/10 · Work/tech: Shopify React Native to native, platforms and SaaS relevant
10th September 2026 - Link Blog Native is now the future of mobile at Shopify (via) Shopify are moving from React Native back to separate Swift and Kotlin codebases for their native apps, for the…

10th September 2026 - Link Blog Native is now the future of mobile at Shopify (via) Shopify are moving from React Native back to separate Swift and Kotlin codebases for their native apps, for the exact reason you would expect: We decided to switch from native to React Native in 2020 for three reasons: - Stop building the same features twice - Allow developers to work across the stack - Spend less time chasing feature parity and more time shipping value [...] Native still means building and maintaining software on two platforms, that cost has not disappeared. What changed is that agents can now do enough of the implementation, translation, testing, and review work that it’s no longer the deciding factor it was in 2020. It's a well-written post, which gives full credit to React Native as a great platform for the six years they were using it. Shopify are the maintainers of three significant React Native libraries: react-native-skia, flash-list, and restyle. The first two are finding new homes; the third "has a smaller user base than our other libraries" and will be archived at the end of 2026. Recent articles - Some thoughts on the Navier–Stokes Millennium Prize Problem - 8th September 2026 - The Pelican comparison grid for Astra is pretty interesting - 4th September 2026 - OpenAI's rogue agents were caught communicating via public wikis - 4th September 2026

9

Quoting Jakub Pachocki

Simon Willison · original → · 7/10 · AI: OpenAI defensive AI strategy, critical AI perspective relevant
7th September 2026 The strongest argument I see for continuing to train much smarter models quickly is the need to build defensive systems against the dangers posed by other AI. [...] We will need…

7th September 2026 The strongest argument I see for continuing to train much smarter models quickly is the need to build defensive systems against the dangers posed by other AI. [...] We will need powerful, aligned AI for defense; to secure infrastructure, to protect against rogue agents in real time, and to invent entirely new protective measures. This will be a primary focus of OpenAI’s deployment efforts. At the same time, even with the uncertainty that comes from anticipated broad AI progress and the need to build defensive systems, we must not let that become an excuse for recklessness. The idea of racing forward at all costs seems absurd once one internalizes the seriousness of the stakes. — Jakub Pachocki, Chief Scientist at OpenAI Recent articles - Some thoughts on the Navier–Stokes Millennium Prize Problem - 8th September 2026 - The Pelican comparison grid for Astra is pretty interesting - 4th September 2026 - OpenAI's rogue agents were caught communicating via public wikis - 4th September 2026

Items scoring 7/10 or above from 11 sources, scored by claude-haiku-4-5-20251001 on relevance to my interests. At most 3 per source.

Scoring categories & sources
  1. Local Wexford or South East Ireland news
  2. Irish or EU-wide affairs affecting citizens broadly: elections, new laws or policy being debated, cost of living, education — especially impacts on mid-life adults or teenagers. Never courts/crime stories.
  3. Irish news on a topic relevant to my interests
  4. Work and tech topics: networking, AI, Kubernetes, platforms, SaaS
  5. AI news including critical or anti-AI perspectives
  6. Gaming: PC gaming, indie gaming, retro gaming
  7. General interests: gardening, woodwork, cycling, fitness, travel
  8. Comics

Sources: Breaking News Ireland, Wexford Local, Hacker News, r/gaming, r/pcgaming, r/antiAI, r/indiegaming, Lenny's Newsletter, One Useful Thing, Newcomer, Simon Willison