Agent engineering
Coding agents, MCP, context systems, evaluation and the controls required for dependable automation.
- Blog posts
- 72
- Topic
- AI and agents
Start with the cornerstone
Get the next AI and agents field note
One concise email when we publish. No tracking pixels, and no inbox filler.
Latest in this collection
AI & AgentsValyu’s 0.6B Multi-Agent Router: Results, Limits and When to Train One
A trained Qwen3-0.6B beat prompted Haiku and Nova Lite in one retrieval-routing study. Understand the metric, training evidence and production decision before adopting it.
AI & AgentsOpenAI Agents API Review: Migration, Costs and Data Controls
What the managed Codex harness replaces, what stays your responsibility, and how to test the launch numbers before migrating a real workflow.
AI & AgentsSpotify shunt Review: Setup, Savings and Limits
What Spotify's 90% claim actually measures, what Portal setup requires, and how to test file-read delegation without moving the cost into rework.
AI & AgentsOpenBot Review: Self-Hosted AI Coworkers, Costs & Controls
A source-reviewed guide to CopilotKit OpenBot: where its action gateway helps, which defaults need changing, and when self-hosting beats buying a managed assistant.
Ramp Inspect Architecture 2026: Background Coding Agents at Scale
How Ramp Inspect uses full-stack Modal sandboxes, filesystem snapshots, coordination and verification, plus the controls a regulated engineering team needs before scaling background coding agents.
Model Hardware Standard: Enterprise Guide to Physical AI
Evaluate Anthropic's MHS for lab and factory automation: architecture, MHS vs MCP, safety boundaries, pilot economics and a practical adoption checklist.
Complete article directory
- Valyu’s 0.6B Multi-Agent Router: Results, Limits and When to Train One
- OpenAI Agents API Review: Migration, Costs and Data Controls
- Spotify shunt Review: Setup, Savings and Limits
- OpenBot Review: Self-Hosted AI Coworkers, Costs & Controls
- Ramp Inspect Architecture 2026: Background Coding Agents at Scale
- Model Hardware Standard: Enterprise Guide to Physical AI
- Fonio AI Review 2026: Pricing, API, GDPR & Build vs Buy
- Mosaic (YC S26) Review: Shared Memory for Team AI Agents
- Ripwire Review 2026: AI Repo Context Without Embeddings?
- Atomic Multi-File Edits for AI Coding Agents: The Semaprax Lesson
- AI Agent Knowledge Transfer: Use Frontier Models Once, Then Scale Cheaper
- claude-rotate: One Proxy for Multiple Claude Max Accounts
- Feynman Review: Is This Open-Source AI Research Agent Ready for Teams?
- Obscura Browser Review: Claims, Limits and Production Fit
- Claude Code Design System: 4 Parts for On-Brand UI
- ChatGPT Can Now Log In Without Seeing Your Password
- AI Agent Harness, Explained: The Reliability Layer Around an LLM
- Why Agent Edits Need Semantic Identity: Building SEMAPRAX in Rust
- LangChain Deep Agents Review: Is the Agent Harness Ready for Production?
- OpenViking Review 2026: Is Filesystem Memory Production-Ready?
- LLM-as-a-Verifier Explained: Architecture, Costs, and Production Fit
- TrueForge Review: Is the Open-Source Agent Harness Production-Ready?
- Agent-Readable Websites: llms.txt, Markdown Mirrors and What Breaks
- Localized URL Slugs vs hreflang: Why We Keep One English Slug
- Can an AI Agent Use Your Product, or Only Read About It?
- Graft Review 2026: Do Agent Repo Maps Belong in Git?
- How Coding Agents Keep Token Bills in Check with Output Compression
- Smarter Token Usage with Your AI Coding Agent
- DeepSeek Harness Review: Is the Plugin Stack Production-Ready?
- OpenSandbox Review: Is Self-Hosting Worth It?
- Cloudflare Kitesurf Review: Cost, Limits and Production Fit
- GitHub Spec Kit Review: Is It Worth the Process?
- Internal AI Agent Marketplace: A 2026 Enterprise Build Guide
- Is Linux the Best OS for AI Agents? A 2026 Infrastructure Guide
- MCP Cloud vs Manufact Cloud: MCP Hosting Guide
- How to Make AI Writing Sound Human with Agent Skills
- NVIDIA NOOA Review: Are Object-Oriented Agents Production-Ready?
- Strix AI Pentesting: 30-Day Pilot and Buying Guide for 2026
- Hyperagent Review: Cloud AI Agents Without a Server
- Hark Handoff Review: The Agent That Actually Clicks
- Meta Muse Code Pricing: Is the Contributor Tier Safe for Client Code?
- PII Redaction Before LLM Prompts: A Practical Pipeline
- Agent Reach Review: Costs, Security and Real Limits
- Cloudflare Wallets for AI Agents: What Is Live?
- YC QM Agent Review: Is Quartermaster Ready for Work?
- jcode vs Claude Code: Is the Rust Harness Worth Switching To?
- Lightpanda Browser for AI Agents: Production Guide
- Graph Engineering for AI Agents: When Does a Knowledge Graph Pay Off?
- Multi-Model AI Coding Agent Stack: A Team Buying Guide
- Is MCP Stateless Now? Your Server Migration Checklist
- MCP Is Not a Security Boundary: Protect Agent Data
- Can AI Agents Talk to Each Other? A Band Setup Guide
- LeanCTX Technical Field Report: 64.1% Less Context
- CLIProxyAPI: Run GPT-5.6 Sol Inside Claude Code
- Enterprise MCP Authorization Architecture
- OpenAI Eval Sandbox Escape: 12 Controls Before You Test a Cyber Agent
- Meterless Review 2026: Is This AI Agent Context Layer Ready?
- Claude Code vs OpenCode: Which Costs Less for a Team in 2026?
- Graphify Review 2026: Is a Codebase Knowledge Graph Worth It?
- AI Pilot Kill-or-Scale Scorecard: 12 Metrics to Check After 30 Days
- T3MP3ST Review 2026: Can It Replace a Penetration Test?
- Open Knowledge Format (OKF): The Enterprise Guide
- AI Agent Cost per Action: Why Agentic Workflows Blow Up Token Bills
- How to Use Claude Fable 5.1 in Claude Code
- What an Internal AI Assistant Actually Costs in the DACH Region (2026)
- The Bottleneck Was Never Intelligence. It Was Context.
- ChatGPT Enterprise vs Copilot vs Custom RAG
- MCP vs RAG vs Agent Skills vs Custom GPTs
- AI Agent Pilot in 30/60/90 Days
- Permissions-First RAG over SharePoint, Confluence, Drive
- RAG Production-Readiness Checklist for EU Companies
- Why AI Agent Projects Get Cancelled