Cluster

Agent engineering

Coding agents, MCP, context systems, evaluation and the controls required for dependable automation.

Blog posts
70
Topic
AI and agents

Start with the cornerstone

Inbox, without the noise

Get the next AI and agents field note

One concise email when we publish. No tracking pixels, and no inbox filler.

Latest in this collection

Large source files go to a worker model while only focused summaries return to Claude; reasoning stays separate AI & Agents

Spotify shunt Review: Setup, Savings and Limits

What Spotify's 90% claim actually measures, what Portal setup requires, and how to test file-read delegation without moving the cost into rework.

Three AI coworker profiles connect to a policy and audit gateway before using browser, file and MCP tools; conversation storage is shown separately AI & Agents

OpenBot Review: Self-Hosted AI Coworkers, Costs & Controls

A source-reviewed guide to CopilotKit OpenBot: where its action gateway helps, which defaults need changing, and when self-hosting beats buying a managed assistant.

A background coding agent receives a prompt, restores a prepared development snapshot, works inside an isolated sandbox, verifies the change and opens a pull request AI & Agents

Ramp Inspect Architecture 2026: Background Coding Agents at Scale

How Ramp Inspect uses full-stack Modal sandboxes, filesystem snapshots, coordination and verification, plus the controls a regulated engineering team needs before scaling background coding agents.

AI agent connecting through a standardized interface to programmable lab and factory hardware AI & Agents

Model Hardware Standard: Enterprise Guide to Physical AI

Evaluate Anthropic's MHS for lab and factory automation: architecture, MHS vs MCP, safety boundaries, pilot economics and a practical adoption checklist.

A voice AI phone platform connected to SIP, business APIs, and a build versus buy decision matrix AI & Agents

Fonio AI Review 2026: Pricing, API, GDPR & Build vs Buy

A buyer-focused Fonio AI review with current phone pricing, SIP and API integration, EU recording duties, operational limits, and a practical build-versus-buy scorecard.

Several coding-agent session timelines converge into one shared team context store AI & Agents

Mosaic (YC S26) Review: Shared Memory for Team AI Agents

Mosaic centralizes coding-agent sessions across tools and teammates. See how shared session memory differs from Claude Code Agent Teams, agent memory, Git, and a real team knowledge layer.

Complete article directory

  1. Spotify shunt Review: Setup, Savings and Limits
  2. OpenBot Review: Self-Hosted AI Coworkers, Costs & Controls
  3. Ramp Inspect Architecture 2026: Background Coding Agents at Scale
  4. Model Hardware Standard: Enterprise Guide to Physical AI
  5. Fonio AI Review 2026: Pricing, API, GDPR & Build vs Buy
  6. Mosaic (YC S26) Review: Shared Memory for Team AI Agents
  7. Ripwire Review 2026: AI Repo Context Without Embeddings?
  8. Atomic Multi-File Edits for AI Coding Agents: The Semaprax Lesson
  9. AI Agent Knowledge Transfer: Use Frontier Models Once, Then Scale Cheaper
  10. claude-rotate: One Proxy for Multiple Claude Max Accounts
  11. Feynman Review: Is This Open-Source AI Research Agent Ready for Teams?
  12. Obscura Browser Review: Claims, Limits and Production Fit
  13. Claude Code Design System: 4 Parts for On-Brand UI
  14. ChatGPT Can Now Log In Without Seeing Your Password
  15. AI Agent Harness, Explained: The Reliability Layer Around an LLM
  16. Why Agent Edits Need Semantic Identity: Building SEMAPRAX in Rust
  17. LangChain Deep Agents Review: Is the Agent Harness Ready for Production?
  18. OpenViking Review 2026: Is Filesystem Memory Production-Ready?
  19. LLM-as-a-Verifier Explained: Architecture, Costs, and Production Fit
  20. TrueForge Review: Is the Open-Source Agent Harness Production-Ready?
  21. Agent-Readable Websites: llms.txt, Markdown Mirrors and What Breaks
  22. Localized URL Slugs vs hreflang: Why We Keep One English Slug
  23. Can an AI Agent Use Your Product, or Only Read About It?
  24. Graft Review 2026: Do Agent Repo Maps Belong in Git?
  25. How Coding Agents Keep Token Bills in Check with Output Compression
  26. Smarter Token Usage with Your AI Coding Agent
  27. DeepSeek Harness Review: Is the Plugin Stack Production-Ready?
  28. OpenSandbox Review: Is Self-Hosting Worth It?
  29. Cloudflare Kitesurf Review: Cost, Limits and Production Fit
  30. GitHub Spec Kit Review: Is It Worth the Process?
  31. Internal AI Agent Marketplace: A 2026 Enterprise Build Guide
  32. Is Linux the Best OS for AI Agents? A 2026 Infrastructure Guide
  33. MCP Cloud vs Manufact Cloud: MCP Hosting Guide
  34. How to Make AI Writing Sound Human with Agent Skills
  35. NVIDIA NOOA Review: Are Object-Oriented Agents Production-Ready?
  36. Strix AI Pentesting: 30-Day Pilot and Buying Guide for 2026
  37. Hyperagent Review: Cloud AI Agents Without a Server
  38. Hark Handoff Review: The Agent That Actually Clicks
  39. Meta Muse Code Pricing: Is the Contributor Tier Safe for Client Code?
  40. PII Redaction Before LLM Prompts: A Practical Pipeline
  41. Agent Reach Review: Costs, Security and Real Limits
  42. Cloudflare Wallets for AI Agents: What Is Live?
  43. YC QM Agent Review: Is Quartermaster Ready for Work?
  44. jcode vs Claude Code: Is the Rust Harness Worth Switching To?
  45. Lightpanda Browser for AI Agents: Production Guide
  46. Graph Engineering for AI Agents: When Does a Knowledge Graph Pay Off?
  47. Multi-Model AI Coding Agent Stack: A Team Buying Guide
  48. Is MCP Stateless Now? Your Server Migration Checklist
  49. MCP Is Not a Security Boundary: Protect Agent Data
  50. Can AI Agents Talk to Each Other? A Band Setup Guide
  51. LeanCTX Technical Field Report: 64.1% Less Context
  52. CLIProxyAPI: Run GPT-5.6 Sol Inside Claude Code
  53. Enterprise MCP Authorization Architecture
  54. OpenAI Eval Sandbox Escape: 12 Controls Before You Test a Cyber Agent
  55. Meterless Review 2026: Is This AI Agent Context Layer Ready?
  56. Claude Code vs OpenCode: Which Costs Less for a Team in 2026?
  57. Graphify Review 2026: Is a Codebase Knowledge Graph Worth It?
  58. AI Pilot Kill-or-Scale Scorecard: 12 Metrics to Check After 30 Days
  59. T3MP3ST Review 2026: Can It Replace a Penetration Test?
  60. Open Knowledge Format (OKF): The Enterprise Guide
  61. AI Agent Cost per Action: Why Agentic Workflows Blow Up Token Bills
  62. How to Use Claude Fable 5.1 in Claude Code
  63. What an Internal AI Assistant Actually Costs in the DACH Region (2026)
  64. The Bottleneck Was Never Intelligence. It Was Context.
  65. ChatGPT Enterprise vs Copilot vs Custom RAG
  66. MCP vs RAG vs Agent Skills vs Custom GPTs
  67. AI Agent Pilot in 30/60/90 Days
  68. Permissions-First RAG over SharePoint, Confluence, Drive
  69. RAG Production-Readiness Checklist for EU Companies
  70. Why AI Agent Projects Get Cancelled