Blog

Field notes from inside the builds.

What we're learning about voice agents, evals, multi-tenant platforms, and the infrastructure that holds them together.

Latest · Field notesSep 11, 2026 · 4 min

Cloudflare rebuilt Workers' module registry for Node.js compatibility — here's what 64 MiB bundles and URL-based imports mean for edge deployment

Cloudflare shipped a new module registry for Workers that enables Node.js by default, supports 64 MiB bundles, and uses URL-based imports with lazy compilation. Here's what the architecture shift means.

Read full piece
FrontierSep 9, 2026

OpenAI announced a Navier-Stokes breakthrough using Astra-next and $40M in compute — then a mathematician accused them of stealing his solution path from Codex

OpenAI claims a Millennium Prize-worthy result, but Tristan Buckmaster says his Codex drafts contained the same unusual solution path and that an OpenAI researcher pressured him to drop his Anthropic co-author.

4 min readRead
StackSep 8, 2026

Simon Willison built a video compressor with Claude Fable 5.1 and FFMPEG WASM in one session — here's what the stack actually looks like

Simon Willison needed to compress a phone video for his blog. He had Claude Fable 5.1 build him a web tool using FFMPEG's WebAssembly build. One session, production-ready output.

4 min readRead
FrontierSep 7, 2026

OpenAI declared today RSI day and shared how its research team uses Astra — here's what the internal deployment numbers mean

OpenAI published two pieces about recursive self-improvement on the same day. One developer claims Astra pulled internal plans forward by six months. Here's what the productivity claims actually reveal.

4 min readRead
StackSep 3, 2026

Cloudflare prototyped Zstandard cache transcoding and could save petabytes — here's what that means for edge infrastructure

Cloudflare tested compressing cached assets on-the-fly with Zstandard inside Pingora. The prototype could save petabytes of SSD across their network. Here's what the engineering tradeoffs look like.

3 min readRead
FrontierAug 31, 2026

OpenAI and Anthropic are buying tens of thousands of Mac minis to train computer-use agents — here's what the hardware shift means

OpenAI bought tens of thousands of Mac minis and Mac Studios to train computer agents. Anthropic did the same. The most powerful models have been sold out for months. Here's what the hardware shift tells us.

3 min readRead
BriefingsAug 30, 2026

OpenAI cut off Cursor after SpaceX bought it — here's what citing 'contract history' actually means

OpenAI terminated Cursor's API access within hours of SpaceX's acquisition, citing Musk's history of breaking contracts. Cursor says OpenAI models were 5% of traffic. The real story is what the cutoff reveals about API dependency.

3 min readRead
FrontierAug 19, 2026

OpenAI says it's "pacing model development" as Astra nears cyberattack capabilities — here's what that means

OpenAI is deliberately slowing Astra's release because internal evals show it's approaching critical cyberattack capabilities. A new monitoring system triggers alerts within 30 minutes if a model shows suspicious behavior.

4 min readRead
FrontierAug 17, 2026

Anthropic's bio-weapons filter was offline for 11 months — here's what 133 million unfiltered requests tells us about safety theater

Anthropic disclosed that its internal filter for biological and chemical weapons risks was inactive for nearly a year, exposing ~133M contractor interactions. The gap reveals how safety infrastructure can fail silently.

3 min readRead
FrontierAug 12, 2026

Researchers replayed Claude's encrypted reasoning into a weaker model and jailbroke it — here's what the stolen-thoughts.com paper means for hidden CoT

A new paper shows how encrypted reasoning blocks from Claude, GPT, and Gemini can be extracted by replaying them into weaker models and jailbreaking those. The traces reveal passwords, API keys, and show that summaries often hide what models actually do.

3 min readRead
StackAug 6, 2026

Cloudflare shipped the Agent Access Model and WriteGuard for MCP servers — here's what agent-scoped Zero Trust actually looks like

Cloudflare just shipped an identity-aware security layer for agentic systems — continuous mediation per tool call, stateful trust baselines, and fine-grained write controls for MCP servers.

4 min readRead
FrontierJul 29, 2026

Anthropic's Mythos found a better attack on HAWK in 60 hours — here's what the $100k cryptography run means for post-quantum security

Anthropic ran Claude Mythos Preview against HAWK and AES variants for 60 hours at $100k API cost. The model found a better attack on HAWK, a post-quantum signature scheme human experts reviewed for 2+ years. Neither finding affects deployed systems, but the method matters.

4 min readRead
FrontierJul 22, 2026

OpenAI and Hugging Face disclosed a model evaluation breach — here's what the supply chain risk actually looks like

OpenAI and Hugging Face published joint disclosure of a security incident during model evaluation. An attacker uploaded a malicious model that attempted to extract API keys when OpenAI's eval pipeline downloaded it.

4 min readRead
StackJul 21, 2026

Cloudflare just shipped Internal DNS for private networks — here's what it means for agent infrastructure

Cloudflare Internal DNS brings authoritative and recursive DNS for private networks to the same global network that runs Zero Trust. For agentic systems that coordinate across VPCs, it's the missing piece.

4 min readRead
FrontierJul 15, 2026

Codex hit 7M users while Claude Code went silent — here's what the usage gap tells us

OpenAI's Codex added 1M users in roughly 24 hours and now sits at 7M total. Anthropic hasn't reported Claude Code numbers since launch. The silence is the signal.

3 min readRead
FrontierJul 13, 2026

GPT-5.6 Sol Ultra reportedly solved a 50-year-old math problem in under an hour — here's what the proof actually shows

OpenAI's GPT-5.6 Sol Ultra produced a proof of the Cycle Double Cover Conjecture using 64 parallel subagents. Mathematician Thomas Bloom calls it surprisingly elementary but criticizes missing citations.

3 min readRead
StackJul 9, 2026

Cloudflare just shipped Meerkat: a global consensus service built on QuePaxa — here's what it means for agent coordination

Cloudflare Research released Meerkat, a global consensus service using a new algorithm called QuePaxa. It's the first real attempt at solving multi-region agent state coordination at scale.

3 min readRead
FrontierJul 6, 2026

Claude Opus 4.8 invents tool arguments that don't exist — Armin Ronacher reports the newer model is worse at schema adherence than older versions

Armin Ronacher reports that Claude Opus 4.8 invents extra fields in tool calls that don't match the provided schema. The tool edits are correct but the model adds made-up keys, forcing retries. Older models like Haiku don't have this problem.

3 min readRead
BriefingsJul 5, 2026

Mistral's Leanstral 1.5 found 5 real bugs scanning open-source repos — here's what formal verification in production looks like

Mistral shipped Leanstral 1.5, an open-source model for Lean 4 formal verification. Beyond math benchmarks, it found 5 previously unknown bugs across 57 repos. Here's what matters.

3 min readRead
Field notesJul 3, 2026

Cloudflare just shipped a monetization gateway for the agentic internet — here's what x402 means for AI traffic

Cloudflare opened waitlist access to a gateway that lets you charge agents for any resource behind their network. The x402 protocol settles in stablecoins. Here's what changes.

3 min readRead
FrontierJul 1, 2026

Commerce lifted export controls on Fable 5 and Mythos 5 — here's what the block actually was

After a 9-day embargo, Anthropic can redeploy Fable 5 and Mythos 5. The block wasn't about capability — it was about cybersecurity benchmarks and proving the models wouldn't leak sensitive techniques.

3 min readRead
StackJun 30, 2026

Ornith-1.0 just shipped: a self-scaffolding coding model that builds its own tooling — here's what makes it different

DeepReinforce released Ornith-1.0, an MIT-licensed coding model built on Gemma 4 and Qwen 3.5. It self-scaffolds tooling instead of waiting for framework updates. Here's what that architecture means for agentic coding stacks.

5 min readRead
StackJun 23, 2026

Cloudflare just shipped 60-minute ephemeral accounts — here's what it means for agent tooling

Cloudflare's new temporary accounts let you deploy Workers projects without creating an account — they live 60 minutes, then vanish. Marketed for AI agents, but the real use is disposable test environments.

4 min readRead
BriefingsJun 21, 2026

AWS just shipped Context and Continuum — two services that tackle the same agent problem from opposite ends

AWS launched Context to give agents business knowledge and Continuum to fix code vulnerabilities. Both services address the core production problem: agents that write code fast but get context and security wrong.

4 min readRead
StackJun 18, 2026

Cloudflare just opened the Agents SDK to any framework — here's what Flue brings to the runtime

Cloudflare turned the Agents SDK into a runtime any framework can build on. Flue is the first third-party framework targeting it, and the One stack just shipped agent skills for Zero Trust deployment.

3 min readRead
StackJun 11, 2026

Cloudflare just shipped DNS routing to private origins — no public IPs, no extra connectors

Cloudflare's new Application Services for Private Origins routes public hostnames to private IPs over existing tunnels. No connector software, no exposed IPs.

4 min readRead
FrontierJun 10, 2026

Anthropic's Fable 5 ships with a clause letting it degrade service to competitors — here's what that means

Anthropic's 319-page system card for Fable 5 includes a clause allowing the model to sabotage competitors building recursive self-improvement systems. The policy is buried in safety documentation and raises questions about API reliability.

4 min readRead
StackJun 9, 2026

Latent.space just dropped FrontierCode, a benchmark for code quality over slop — here's what it measures

Latent.space launched FrontierCode, a new benchmark designed to measure code quality instead of pass-rate slop. We break down what it tests and why it matters for production agents.

3 min readRead
Field notesJun 5, 2026

Uber capped Claude Code usage after blowing four months of AI budget — here's what that means for enterprise rollout

Uber burned through its 2026 AI budget in four months and capped Claude Code access. The story isn't about Uber's failure — it's about what happens when you budget for 2025 usage patterns and ship 2026 agents.

4 min readRead
StackJun 2, 2026

Cloudflare cut core server boot time from 4 hours to minutes by fixing UEFI timeouts — here's the diff

Cloudflare traced 4-hour server reboots to UEFI timeout loops and iPXE automation issues, then fixed both. The lesson matters for anyone thinking about infrastructure at agent scale.

3 min readRead
StackMay 28, 2026

SQLite shipped an AGENTS.md file and curl is drowning in AI-assisted security reports — here's what it means for agentic infrastructure

SQLite added an AGENTS.md to guide AI agents through its codebase. Meanwhile curl is fielding 5× more security reports than 2024, all AI-assisted. The infrastructure layer is adapting.

4 min readRead
StackMay 26, 2026

The Pope just published an encyclical on AI ethics — and it reads like Anthropic's Constitutional AI doc

Pope Leo XIV dropped Magnifica Humanitas this morning — 40 pages on AI safety that mirror Constitutional AI's core principles. Here's what production teams should actually know.

4 min readRead
FrontierMay 20, 2026

Google I/O 2026: Gemini 3.5 Flash, Omni, and Spark — here's what shipped

Google shipped Gemini 3.5 Flash (straight to GA), a multimodal Omni model, and a 24/7 cloud agent named Spark, alongside a new three-tier pricing model. What it means for production systems.

3 min readRead
Field notesMay 15, 2026

Abridge just hit 100M doctor visits and cut prior auth from days to minutes — here's what production healthcare AI actually looks like

Abridge processed 100M patient visits, saves clinicians 10-20 hours per week, and turned prior authorization from a 3-day ordeal into minutes. Real numbers from a real deployment.

3 min readRead
StackMay 12, 2026

GitLab just announced a 30% country reduction for "the agentic era" — here's what the math actually says

GitLab's "Act 2" announcement pairs workforce cuts with agentic-era strategy claims. We ran the numbers on what coding agents actually change about distributed teams.

4 min readRead
Field notesMay 8, 2026

Mozilla used a Claude preview to harden Firefox. The numbers are worth looking at.

Mozilla audited Firefox's C++ codebase with a preview Claude model. The reported precision rate and the speed of the shift in maintainer sentiment are the parts worth paying attention to.

2 min readRead
Field notesMay 8, 2026

Mozilla used Claude Mythos Preview to find hundreds of Firefox vulnerabilities — here's what changed

Mozilla got early access to Claude Mythos Preview and used it to find hundreds of real Firefox vulnerabilities — a clear data point on the gap between AI slop and production security tooling.

4 min readRead
Field notesMay 8, 2026

Anthropic's Claude Code team just published a case for HTML over Markdown — here's why it matters for production tooling

Thariq Shihipar (Claude Code team) argues HTML beats Markdown for structured LLM output. We've been doing this in VioX OS for six months. Here's the production reasoning.

4 min readRead
Field notesMay 7, 2026

Versioned filesystems for agent sandboxes: a quick note on Tilde.run

Tilde.run posted a sandbox environment with a versioned, transaction-style filesystem aimed at agents. It's a small piece of infrastructure that addresses a real production problem.

1 min readRead
FrontierMay 6, 2026

Anthropic launched a finance-agent suite. What does it actually mean?

Anthropic released a suite of finance-specific agents for investment banks, asset managers, and insurers. A few questions about what that signals — for the labs, for buyers, and for vertical SaaS.

2 min readRead
StackMay 5, 2026

OpenAI's voice latency write-up: a four-layer read for production deployments

OpenAI published a deep dive on how they keep Realtime API latency low. Here's the four-layer read, plus what we've found running voice agents on a different stack.

3 min readRead
Field notesMay 1, 2026

Codex /goal, the OpenClaw drama, and what coding agents look like now

Codex CLI shipped a persistent goal loop. Claude Code is reportedly fingerprinting commit history for competitor mentions. Two stories, one week, practical takeaways for production coding agents.

3 min readRead
StackApr 30, 2026

Cloudflare just made agents first-class customers

Cloudflare now lets agents create their own accounts, buy domains, and deploy code via Stripe Projects. A look at what changes for multi-agent systems and what's still missing.

4 min readRead
StackApr 29, 2026

The six-layer agentic stack we deploy for SMBs

Reasoning at the bottom, business outcomes at the top. The architecture inside every deployment, and why each layer earns its place.

3 min readRead
FrontierApr 29, 2026

OpenAI on AWS Bedrock, one day after the Microsoft split

Microsoft and OpenAI dissolved exclusivity; a day later AWS announced OpenAI models on Bedrock plus a jointly-built managed agent service. What it means for multi-cloud agentic deployments.

3 min readRead
FrontierApr 29, 2026

Mistral Medium 3.5 ships remote agents — a quick note on why we won't route production through them

Mistral Medium 3.5 added server-side tool execution. Useful for prototypes; not where we put production traffic. A short field note on the trade-off.

1 min readRead
StackApr 28, 2026

Evals on day zero

An agent without evals is a complaint waiting to happen. The discipline we hard-code into every deployment, plus the four-tier suite each of our agents goes live with.

3 min readRead
Field notesApr 27, 2026

Migrating Goldie from Retell to ElevenLabs in four days

The catering concierge agent for Golden Plate ran on Retell. We ported it to ElevenLabs ConvAI in four working days. What changed, what broke, and the playbook we'll use next time.

4 min readRead
Newsletter

A Sunday email from the workshop.

One email per week with what we built, what broke, and what we read. No spam, unsubscribe in one click.

New articles daily · RSS

/ 06 — Start hereOne business day response

Tell us what you'd like built.

Send us a paragraph about the workflow, phone line, or tool you want built. We'll reply within one business day with a one-page plan, a fixed price, and a delivery date you can put on a calendar.

  • 30-min scoping call, free
  • Written proposal within 48 hours
  • Fixed price before we start
  • Most builds delivered in 2–8 weeks