Agent engineering
Coding agents, MCP, context systems, evaluation and the controls required for dependable automation.
- Blog posts
- 55
- Topic
- AI and agents
Start with the cornerstone
Latest in this collection
Why Agent Edits Need Semantic Identity: Building SEMAPRAX in Rust
A compiler engineering walkthrough of persistent IDs, checked HIR, deterministic graphs and replayable patches in the experimental SEMAPRAX language.
LangChain Deep Agents Review: Is the Agent Harness Ready for Production?
A buyer-focused Deep Agents review covering architecture, v0.7 changes, security boundaries, operating costs and a measurable production pilot.
OpenViking Review 2026: Is Filesystem Memory Production-Ready?
Buyer review of OpenViking for agent memory and RAG: architecture, AGPL risk, security controls, operating cost and a measurable two-week pilot.
LLM-as-a-Verifier Explained: Architecture, Costs, and Production Fit
A technical buyer guide to fine-grained LLM verification: how logprob scoring and pivot tournaments work, what the benchmarks prove, and how to pilot it safely.
TrueForge Review: Is the Open-Source Agent Harness Production-Ready?
A buyer-focused review of TrueForge, its agent loop, context controls, benchmark economics, deployment modes and enterprise production gaps.
Agent-Readable Websites: llms.txt, Markdown Mirrors and What Breaks
Four surfaces decide whether an AI answer engine can read you, and every failure mode is invisible in a browser. Here is the list our own build gate has caught.
Complete article directory
- Why Agent Edits Need Semantic Identity: Building SEMAPRAX in Rust
- LangChain Deep Agents Review: Is the Agent Harness Ready for Production?
- OpenViking Review 2026: Is Filesystem Memory Production-Ready?
- LLM-as-a-Verifier Explained: Architecture, Costs, and Production Fit
- TrueForge Review: Is the Open-Source Agent Harness Production-Ready?
- Agent-Readable Websites: llms.txt, Markdown Mirrors and What Breaks
- Localized URLs Break hreflang: Keep One English Slug
- Can an AI Agent Use Your Product, or Only Read About It?
- Graft Review 2026: Do Agent Repo Maps Belong in Git?
- How Coding Agents Keep Token Bills in Check with Output Compression
- Smarter Token Usage with Your AI Coding Agent
- DeepSeek Harness Review: Is the Plugin Stack Production-Ready?
- OpenSandbox Review: Is Self-Hosting Worth It?
- Cloudflare Kitesurf Review: Cost, Limits and Production Fit
- GitHub Spec Kit Review: Is It Worth the Process?
- Internal AI Agent Marketplace: A 2026 Enterprise Build Guide
- Is Linux the Best OS for AI Agents? A 2026 Infrastructure Guide
- MCP Cloud vs Manufact Cloud: MCP Hosting Guide
- How to Make AI Writing Sound Human with Agent Skills
- NVIDIA NOOA Review: Are Object-Oriented Agents Production-Ready?
- Strix AI Pentesting: 30-Day Pilot and Buying Guide for 2026
- Hyperagent Review: Cloud AI Agents Without a Server
- Hark Handoff Review: The Agent That Actually Clicks
- Meta Muse Code Pricing: Is the Contributor Tier Safe for Client Code?
- PII Redaction Before LLM Prompts: A Practical Pipeline
- Agent Reach Review: Costs, Security and Real Limits
- Cloudflare Wallets for AI Agents: What Is Live?
- YC QM Agent Review: Is Quartermaster Ready for Work?
- jcode vs Claude Code: Is the Rust Harness Worth Switching To?
- Lightpanda Browser for AI Agents: Production Guide
- Graph Engineering for AI Agents: When Does a Knowledge Graph Pay Off?
- Multi-Model AI Coding Agent Stack: A Team Buying Guide
- Is MCP Stateless Now? Your Server Migration Checklist
- MCP Is Not a Security Boundary: Protect Agent Data
- Can AI Agents Talk to Each Other? A Band Setup Guide
- LeanCTX Technical Field Report: 64.1% Less Context
- CLIProxyAPI: Run GPT-5.6 Sol Inside Claude Code
- Enterprise MCP Authorization Architecture
- OpenAI Eval Sandbox Escape: 12 Controls Before You Test a Cyber Agent
- Meterless Review 2026: Is This AI Agent Context Layer Ready?
- Claude Code vs OpenCode: Which Costs Less for a Team in 2026?
- Graphify Review 2026: Is a Codebase Knowledge Graph Worth It?
- AI Pilot Kill-or-Scale Scorecard: 12 Metrics to Check After 30 Days
- T3MP3ST Review 2026: Can It Replace a Penetration Test?
- Open Knowledge Format (OKF): The Enterprise Guide
- AI Agent Cost per Action: Why Agentic Workflows Blow Up Token Bills
- How to Use Claude Fable 5 in Claude Code
- What an Internal AI Assistant Actually Costs in the DACH Region (2026)
- The Bottleneck Was Never Intelligence. It Was Context.
- ChatGPT Enterprise vs Copilot vs Custom RAG
- MCP vs RAG vs Agent Skills vs Custom GPTs
- AI Agent Pilot in 30/60/90 Days
- Permissions-First RAG over SharePoint, Confluence, Drive
- RAG Production-Readiness Checklist for EU Companies
- Why 40% of AI Agent Projects Die