Cluster

Agent engineering

Coding agents, MCP, context systems, evaluation and the controls required for dependable automation.

Blog posts
72
Topic
AI and agents

Start with the cornerstone

Inbox, without the noise

Get the next AI and agents field note

One concise email when we publish. No tracking pixels, and no inbox filler.

Latest in this collection

Teal routing node branching to three endpoints on a dark Wavect header, with one selected route highlighted AI & Agents

Valyu’s 0.6B Multi-Agent Router: Results, Limits and When to Train One

A trained Qwen3-0.6B beat prompted Haiku and Nova Lite in one retrieval-routing study. Understand the metric, training evidence and production decision before adopting it.

Geometric code core and four tool nodes inside layered runtime frames on a dark Wavect header AI & Agents

OpenAI Agents API Review: Migration, Costs and Data Controls

What the managed Codex harness replaces, what stays your responsibility, and how to test the launch numbers before migrating a real workflow.

Large source files go to a worker model while only focused summaries return to Claude; reasoning stays separate AI & Agents

Spotify shunt Review: Setup, Savings and Limits

What Spotify's 90% claim actually measures, what Portal setup requires, and how to test file-read delegation without moving the cost into rework.

Three AI coworker profiles connect to a policy and audit gateway before using browser, file and MCP tools; conversation storage is shown separately AI & Agents

OpenBot Review: Self-Hosted AI Coworkers, Costs & Controls

A source-reviewed guide to CopilotKit OpenBot: where its action gateway helps, which defaults need changing, and when self-hosting beats buying a managed assistant.

A background coding agent receives a prompt, restores a prepared development snapshot, works inside an isolated sandbox, verifies the change and opens a pull request AI & Agents

Ramp Inspect Architecture 2026: Background Coding Agents at Scale

How Ramp Inspect uses full-stack Modal sandboxes, filesystem snapshots, coordination and verification, plus the controls a regulated engineering team needs before scaling background coding agents.

AI agent connecting through a standardized interface to programmable lab and factory hardware AI & Agents

Model Hardware Standard: Enterprise Guide to Physical AI

Evaluate Anthropic's MHS for lab and factory automation: architecture, MHS vs MCP, safety boundaries, pilot economics and a practical adoption checklist.

Complete article directory

  1. Valyu’s 0.6B Multi-Agent Router: Results, Limits and When to Train One
  2. OpenAI Agents API Review: Migration, Costs and Data Controls
  3. Spotify shunt Review: Setup, Savings and Limits
  4. OpenBot Review: Self-Hosted AI Coworkers, Costs & Controls
  5. Ramp Inspect Architecture 2026: Background Coding Agents at Scale
  6. Model Hardware Standard: Enterprise Guide to Physical AI
  7. Fonio AI Review 2026: Pricing, API, GDPR & Build vs Buy
  8. Mosaic (YC S26) Review: Shared Memory for Team AI Agents
  9. Ripwire Review 2026: AI Repo Context Without Embeddings?
  10. Atomic Multi-File Edits for AI Coding Agents: The Semaprax Lesson
  11. AI Agent Knowledge Transfer: Use Frontier Models Once, Then Scale Cheaper
  12. claude-rotate: One Proxy for Multiple Claude Max Accounts
  13. Feynman Review: Is This Open-Source AI Research Agent Ready for Teams?
  14. Obscura Browser Review: Claims, Limits and Production Fit
  15. Claude Code Design System: 4 Parts for On-Brand UI
  16. ChatGPT Can Now Log In Without Seeing Your Password
  17. AI Agent Harness, Explained: The Reliability Layer Around an LLM
  18. Why Agent Edits Need Semantic Identity: Building SEMAPRAX in Rust
  19. LangChain Deep Agents Review: Is the Agent Harness Ready for Production?
  20. OpenViking Review 2026: Is Filesystem Memory Production-Ready?
  21. LLM-as-a-Verifier Explained: Architecture, Costs, and Production Fit
  22. TrueForge Review: Is the Open-Source Agent Harness Production-Ready?
  23. Agent-Readable Websites: llms.txt, Markdown Mirrors and What Breaks
  24. Localized URL Slugs vs hreflang: Why We Keep One English Slug
  25. Can an AI Agent Use Your Product, or Only Read About It?
  26. Graft Review 2026: Do Agent Repo Maps Belong in Git?
  27. How Coding Agents Keep Token Bills in Check with Output Compression
  28. Smarter Token Usage with Your AI Coding Agent
  29. DeepSeek Harness Review: Is the Plugin Stack Production-Ready?
  30. OpenSandbox Review: Is Self-Hosting Worth It?
  31. Cloudflare Kitesurf Review: Cost, Limits and Production Fit
  32. GitHub Spec Kit Review: Is It Worth the Process?
  33. Internal AI Agent Marketplace: A 2026 Enterprise Build Guide
  34. Is Linux the Best OS for AI Agents? A 2026 Infrastructure Guide
  35. MCP Cloud vs Manufact Cloud: MCP Hosting Guide
  36. How to Make AI Writing Sound Human with Agent Skills
  37. NVIDIA NOOA Review: Are Object-Oriented Agents Production-Ready?
  38. Strix AI Pentesting: 30-Day Pilot and Buying Guide for 2026
  39. Hyperagent Review: Cloud AI Agents Without a Server
  40. Hark Handoff Review: The Agent That Actually Clicks
  41. Meta Muse Code Pricing: Is the Contributor Tier Safe for Client Code?
  42. PII Redaction Before LLM Prompts: A Practical Pipeline
  43. Agent Reach Review: Costs, Security and Real Limits
  44. Cloudflare Wallets for AI Agents: What Is Live?
  45. YC QM Agent Review: Is Quartermaster Ready for Work?
  46. jcode vs Claude Code: Is the Rust Harness Worth Switching To?
  47. Lightpanda Browser for AI Agents: Production Guide
  48. Graph Engineering for AI Agents: When Does a Knowledge Graph Pay Off?
  49. Multi-Model AI Coding Agent Stack: A Team Buying Guide
  50. Is MCP Stateless Now? Your Server Migration Checklist
  51. MCP Is Not a Security Boundary: Protect Agent Data
  52. Can AI Agents Talk to Each Other? A Band Setup Guide
  53. LeanCTX Technical Field Report: 64.1% Less Context
  54. CLIProxyAPI: Run GPT-5.6 Sol Inside Claude Code
  55. Enterprise MCP Authorization Architecture
  56. OpenAI Eval Sandbox Escape: 12 Controls Before You Test a Cyber Agent
  57. Meterless Review 2026: Is This AI Agent Context Layer Ready?
  58. Claude Code vs OpenCode: Which Costs Less for a Team in 2026?
  59. Graphify Review 2026: Is a Codebase Knowledge Graph Worth It?
  60. AI Pilot Kill-or-Scale Scorecard: 12 Metrics to Check After 30 Days
  61. T3MP3ST Review 2026: Can It Replace a Penetration Test?
  62. Open Knowledge Format (OKF): The Enterprise Guide
  63. AI Agent Cost per Action: Why Agentic Workflows Blow Up Token Bills
  64. How to Use Claude Fable 5.1 in Claude Code
  65. What an Internal AI Assistant Actually Costs in the DACH Region (2026)
  66. The Bottleneck Was Never Intelligence. It Was Context.
  67. ChatGPT Enterprise vs Copilot vs Custom RAG
  68. MCP vs RAG vs Agent Skills vs Custom GPTs
  69. AI Agent Pilot in 30/60/90 Days
  70. Permissions-First RAG over SharePoint, Confluence, Drive
  71. RAG Production-Readiness Checklist for EU Companies
  72. Why AI Agent Projects Get Cancelled