Topic

AI and agents

Engineering, operating and evaluating AI systems, coding agents and model infrastructure.

Blog posts
127
Focused clusters
02

Focused clusters

Start with the cornerstone

Inbox, without the noise

Get the next AI and agents field note

One concise email when we publish. No tracking pixels, and no inbox filler.

Latest in this collection

Teal routing node branching to three endpoints on a dark Wavect header, with one selected route highlighted AI & Agents

Valyu’s 0.6B Multi-Agent Router: Results, Limits and When to Train One

A trained Qwen3-0.6B beat prompted Haiku and Nova Lite in one retrieval-routing study. Understand the metric, training evidence and production decision before adopting it.

Geometric code core and four tool nodes inside layered runtime frames on a dark Wavect header AI & Agents

OpenAI Agents API Review: Migration, Costs and Data Controls

What the managed Codex harness replaces, what stays your responsibility, and how to test the launch numbers before migrating a real workflow.

Large source files go to a worker model while only focused summaries return to Claude; reasoning stays separate AI & Agents

Spotify shunt Review: Setup, Savings and Limits

What Spotify's 90% claim actually measures, what Portal setup requires, and how to test file-read delegation without moving the cost into rework.

Three AI coworker profiles connect to a policy and audit gateway before using browser, file and MCP tools; conversation storage is shown separately AI & Agents

OpenBot Review: Self-Hosted AI Coworkers, Costs & Controls

A source-reviewed guide to CopilotKit OpenBot: where its action gateway helps, which defaults need changing, and when self-hosting beats buying a managed assistant.

A phone and laptop split one large language model across browser tabs while WebRTC carries activations between them AI & Agents

SwarmLLM Review 2026: Browser P2P LLM Inference Across Phones and Laptops

A buyer-focused review of SwarmLLM: browser-native 27B inference across phones and laptops, WebGPU/WebRTC architecture, benchmark limits, peer trust, and when to pilot it.

A background coding agent receives a prompt, restores a prepared development snapshot, works inside an isolated sandbox, verifies the change and opens a pull request AI & Agents

Ramp Inspect Architecture 2026: Background Coding Agents at Scale

How Ramp Inspect uses full-stack Modal sandboxes, filesystem snapshots, coordination and verification, plus the controls a regulated engineering team needs before scaling background coding agents.

Complete article directory

  1. Valyu’s 0.6B Multi-Agent Router: Results, Limits and When to Train One
  2. OpenAI Agents API Review: Migration, Costs and Data Controls
  3. Spotify shunt Review: Setup, Savings and Limits
  4. OpenBot Review: Self-Hosted AI Coworkers, Costs & Controls
  5. SwarmLLM Review 2026: Browser P2P LLM Inference Across Phones and Laptops
  6. Ramp Inspect Architecture 2026: Background Coding Agents at Scale
  7. Model Hardware Standard: Enterprise Guide to Physical AI
  8. Fonio AI Review 2026: Pricing, API, GDPR & Build vs Buy
  9. Mosaic (YC S26) Review: Shared Memory for Team AI Agents
  10. Ripwire Review 2026: AI Repo Context Without Embeddings?
  11. Atomic Multi-File Edits for AI Coding Agents: The Semaprax Lesson
  12. Phonely Alma Review: Is the Voice LLM Ready for Production?
  13. Utopia Review: Temporal Knowledge Graph for Enterprise
  14. AI Agent Knowledge Transfer: Use Frontier Models Once, Then Scale Cheaper
  15. claude-rotate: One Proxy for Multiple Claude Max Accounts
  16. NVIDIA PAIR Review: Local AI Routing and the AMD Gap
  17. Feynman Review: Is This Open-Source AI Research Agent Ready for Teams?
  18. Obscura Browser Review: Claims, Limits and Production Fit
  19. Claude Code Design System: 4 Parts for On-Brand UI
  20. Tencent Hy4 Preview Review: Is the 1M-Context Coding Model Worth a Pilot?
  21. PageLM Review: Self-Hosting and Commercial Use
  22. ChatGPT Can Now Log In Without Seeing Your Password
  23. M6 Mac mini vs M5 Mac Studio for Local AI: Buyer's Guide
  24. AI Agent Harness, Explained: The Reliability Layer Around an LLM
  25. FreeToken AI Review: Run Frontier MoE Models on Consumer GPUs?
  26. Darkbloom AI Review: Private Inference on Idle Macs
  27. Ox Alpha Revealed as GLM-5.3-Flash: What Changed
  28. Why Agent Edits Need Semantic Identity: Building SEMAPRAX in Rust
  29. LangChain Deep Agents Review: Is the Agent Harness Ready for Production?
  30. Pika Audio Models API Pricing: Is SFX Really 20x Cheaper?
  31. Thunder Compute's $13M GPU Virtualization Bet: Enterprise Buyer's Guide
  32. OpenViking Review 2026: Is Filesystem Memory Production-Ready?
  33. LLM-as-a-Verifier Explained: Architecture, Costs, and Production Fit
  34. AirLLM on 4 GB VRAM: How Layer-Wise Inference Really Works
  35. TrueForge Review: Is the Open-Source Agent Harness Production-Ready?
  36. Agent-Readable Websites: llms.txt, Markdown Mirrors and What Breaks
  37. Qwen3.8-27B: Self-Hosted Computer-Use Agents Without Exporting Screenshots
  38. Localized URL Slugs vs hreflang: Why We Keep One English Slug
  39. Can an AI Agent Use Your Product, or Only Read About It?
  40. Graft Review 2026: Do Agent Repo Maps Belong in Git?
  41. How Coding Agents Keep Token Bills in Check with Output Compression
  42. Smarter Token Usage with Your AI Coding Agent
  43. DeepSeek Harness Review: Is the Plugin Stack Production-Ready?
  44. Netflix's vLLM and Triton Stack: 7 Production Lessons
  45. OpenSandbox Review: Is Self-Hosting Worth It?
  46. Transformers.js Browser AI: When Local Inference Belongs in Your Product
  47. Cloudflare Kitesurf Review: Cost, Limits and Production Fit
  48. GitHub Spec Kit Review: Is It Worth the Process?
  49. Internal AI Agent Marketplace: A 2026 Enterprise Build Guide
  50. How to Self-Host LiteLLM in Production: 2026 Guide
  51. AI-Ready Company Wiki: Architecture and Build Guide
  52. Is Linux the Best OS for AI Agents? A 2026 Infrastructure Guide
  53. Does Claude Watermark Text? The 2026 API Answer
  54. OpenKB Review: Knowledge Compiler vs RAG
  55. Unsloth Desktop Review: A Private Local AI Workstation?
  56. NeMo Switchyard 0.2: Agent Model Routing Without Training?
  57. Firecrawl AnyDoc Review: 14 Formats to Markdown
  58. MCP Cloud vs Manufact Cloud: MCP Hosting Guide
  59. Muse Glimmer 30B: Is Meta's Local Agent Model Production-Ready?
  60. How to Make AI Writing Sound Human with Agent Skills
  61. NVIDIA NOOA Review: Are Object-Oriented Agents Production-Ready?
  62. OmniRoute AI Routing: Setup and Production Checklist
  63. Strix AI Pentesting: 30-Day Pilot and Buying Guide for 2026
  64. Hyperagent Review: Cloud AI Agents Without a Server
  65. Gemini Robotics 2: Whole-Body Control and the Pilot Decision
  66. Hark Handoff Review: The Agent That Actually Clicks
  67. Meta Muse Code Pricing: Is the Contributor Tier Safe for Client Code?
  68. pdf-inspector Review: Route PDFs Before OCR
  69. PII Redaction Before LLM Prompts: A Practical Pipeline
  70. Agent Reach Review: Costs, Security and Real Limits
  71. Cloudflare Wallets for AI Agents: What Is Live?
  72. Local Multimodal AI Coding Assistant: Voice, OCR and Privacy
  73. DeepSeek V4 Flash 0731 on One AI PC: What Actually Works?
  74. YC QM Agent Review: Is Quartermaster Ready for Work?
  75. jcode vs Claude Code: Is the Rust Harness Worth Switching To?
  76. Lightpanda Browser for AI Agents: Production Guide
  77. Graph Engineering for AI Agents: When Does a Knowledge Graph Pay Off?
  78. Multi-Model AI Coding Agent Stack: A Team Buying Guide
  79. Is MCP Stateless Now? Your Server Migration Checklist
  80. MCP Is Not a Security Boundary: Protect Agent Data
  81. Fine-Tune Gemma 4 Free with Unsloth and Colab
  82. Can AI Agents Talk to Each Other? A Band Setup Guide
  83. Taalas HC1 Review: Is a Hardwired LLM ASIC Worth It?
  84. LeanCTX Technical Field Report: 64.1% Less Context
  85. llmfit Guide: Which Local LLM Fits Your Hardware?
  86. CLIProxyAPI: Run GPT-5.6 Sol Inside Claude Code
  87. Miso TTS Self-Hosted vs API: Cost, Latency and VRAM Reality
  88. Enterprise MCP Authorization Architecture
  89. Compress RAG Vectors to 2 Bits: Is Data-Oblivious Quantization Production-Ready?
  90. LMCache Reported 13.7x Lower Mean TTFT with a Shared KV Cache
  91. Lossless LLM Weight Compression vs 8-bit GGUF: What Is Ready for Production?
  92. OpenAI Eval Sandbox Escape: 12 Controls Before You Test a Cyber Agent
  93. Cisco Antares Review: 1B Local AI for Vulnerability Localization
  94. Meterless Review 2026: Is This AI Agent Context Layer Ready?
  95. NVIDIA Nemotron 3.5 ASR: Is Free Self-Hosted STT Ready for Voice Agents?
  96. Claude Code vs OpenCode: Which Costs Less for a Team in 2026?
  97. Kimi K3 for EU Companies: API Cost, Data Risk, and a Pilot Plan
  98. Mesh LLM Review: Can One Large LLM Run Across Multiple Computers?
  99. Graphify Review 2026: Is a Codebase Knowledge Graph Worth It?
  100. Bonsai 27B Review: Can a 27B LLM Really Run on a Phone?
  101. Soofi S: Is Germany's Sovereign LLM Ready for Business?
  102. AI Pilot Kill-or-Scale Scorecard: 12 Metrics to Check After 30 Days
  103. T3MP3ST Review 2026: Can It Replace a Penetration Test?
  104. Colibri Runs GLM-5.2 on Consumer Hardware. Here Is the Catch.
  105. Open Knowledge Format (OKF): The Enterprise Guide
  106. AI Agent Cost per Action: Why Agentic Workflows Blow Up Token Bills
  107. When Local Models Beat APIs: A Break-Even Calculator for EU Companies
  108. LLM Cost Calculator 2026: Cost per Task, Not Cost per Token
  109. Pxpipe Review: Can Images Cut Claude Code Token Costs 60%?
  110. Cheaper Per Token Can Still Cost More Per Task
  111. How to Use Claude Fable 5.1 in Claude Code
  112. Anthropic Says No Open-Weights Ban. The Cost War Is Still Real.
  113. What an Internal AI Assistant Actually Costs in the DACH Region (2026)
  114. LLM Gateways Compared 2026: LiteLLM vs OpenRouter vs Portkey vs RouteLLM
  115. Self-Hosting LLMs in the EU: When Open Weights Actually Pay Off
  116. Best Open-Weight LLMs 2026: DeepSeek vs Qwen vs Kimi vs GLM vs Llama
  117. The Bottleneck Was Never Intelligence. It Was Context.
  118. ChatGPT Enterprise vs Copilot vs Custom RAG
  119. MCP vs RAG vs Agent Skills vs Custom GPTs
  120. How to Cut LLM Token Costs in 2026
  121. AI Agent Pilot in 30/60/90 Days
  122. Permissions-First RAG over SharePoint, Confluence, Drive
  123. RAG Production-Readiness Checklist for EU Companies
  124. When Is an LLM Eval Worth Building? Cost, ROI, and Trusting the Judge
  125. Why AI Agent Projects Get Cancelled
  126. RAG vs Fine-Tuning vs Long-Context 2026
  127. LLM API Costs 2026. Architecture Shift