Topic

AI and agents

Engineering, operating and evaluating AI systems, coding agents and model infrastructure.

Blog posts
128
Focused clusters
02

Focused clusters

Start with the cornerstone

Inbox, without the noise

Get the next AI and agents field note

One concise email when we publish. No tracking pixels, and no inbox filler.

Latest in this collection

Wavect editorial illustration for Voicebox local voice cloning and MCP integration AI & Agents

Voicebox: Local Voice Cloning, Dictation and MCP Setup

Clone a voice, dictate across apps and let an existing agent speak locally. A sourced Voicebox guide with MCP and REST setup, engine choices and privacy checks.

Teal routing node branching to three endpoints on a dark Wavect header, with one selected route highlighted AI & Agents

Valyu’s 0.6B Multi-Agent Router: Results, Limits and When to Train One

A trained Qwen3-0.6B beat prompted Haiku and Nova Lite in one retrieval-routing study. Understand the metric, training evidence and production decision before adopting it.

Geometric code core and four tool nodes inside layered runtime frames on a dark Wavect header AI & Agents

OpenAI Agents API Review: Migration, Costs and Data Controls

What the managed Codex harness replaces, what stays your responsibility, and how to test the launch numbers before migrating a real workflow.

Large source files go to a worker model while only focused summaries return to Claude; reasoning stays separate AI & Agents

Spotify shunt Review: Setup, Savings and Limits

What Spotify's 90% claim actually measures, what Portal setup requires, and how to test file-read delegation without moving the cost into rework.

Three AI coworker profiles connect to a policy and audit gateway before using browser, file and MCP tools; conversation storage is shown separately AI & Agents

OpenBot Review: Self-Hosted AI Coworkers, Costs & Controls

A source-reviewed guide to CopilotKit OpenBot: where its action gateway helps, which defaults need changing, and when self-hosting beats buying a managed assistant.

A phone and laptop split one large language model across browser tabs while WebRTC carries activations between them AI & Agents

SwarmLLM Review 2026: Browser P2P LLM Inference Across Phones and Laptops

A buyer-focused review of SwarmLLM: browser-native 27B inference across phones and laptops, WebGPU/WebRTC architecture, benchmark limits, peer trust, and when to pilot it.

Complete article directory

  1. Voicebox: Local Voice Cloning, Dictation and MCP Setup
  2. Valyu’s 0.6B Multi-Agent Router: Results, Limits and When to Train One
  3. OpenAI Agents API Review: Migration, Costs and Data Controls
  4. Spotify shunt Review: Setup, Savings and Limits
  5. OpenBot Review: Self-Hosted AI Coworkers, Costs & Controls
  6. SwarmLLM Review 2026: Browser P2P LLM Inference Across Phones and Laptops
  7. Ramp Inspect Architecture 2026: Background Coding Agents at Scale
  8. Model Hardware Standard: Enterprise Guide to Physical AI
  9. Fonio AI Review 2026: Pricing, API, GDPR & Build vs Buy
  10. Mosaic (YC S26) Review: Shared Memory for Team AI Agents
  11. Ripwire Review 2026: AI Repo Context Without Embeddings?
  12. Atomic Multi-File Edits for AI Coding Agents: The Semaprax Lesson
  13. Phonely Alma Review: Is the Voice LLM Ready for Production?
  14. Utopia Review: Temporal Knowledge Graph for Enterprise
  15. AI Agent Knowledge Transfer: Use Frontier Models Once, Then Scale Cheaper
  16. claude-rotate: One Proxy for Multiple Claude Max Accounts
  17. NVIDIA PAIR Review: Local AI Routing and the AMD Gap
  18. Feynman Review: Is This Open-Source AI Research Agent Ready for Teams?
  19. Obscura Browser Review: Claims, Limits and Production Fit
  20. Claude Code Design System: 4 Parts for On-Brand UI
  21. Tencent Hy4 Preview Review: Is the 1M-Context Coding Model Worth a Pilot?
  22. PageLM Review: Self-Hosting and Commercial Use
  23. ChatGPT Can Now Log In Without Seeing Your Password
  24. M6 Mac mini vs M5 Mac Studio for Local AI: Buyer's Guide
  25. AI Agent Harness, Explained: The Reliability Layer Around an LLM
  26. FreeToken AI Review: Run Frontier MoE Models on Consumer GPUs?
  27. Darkbloom AI Review: Private Inference on Idle Macs
  28. Ox Alpha Revealed as GLM-5.3-Flash: What Changed
  29. Why Agent Edits Need Semantic Identity: Building SEMAPRAX in Rust
  30. LangChain Deep Agents Review: Is the Agent Harness Ready for Production?
  31. Pika Audio Models API Pricing: Is SFX Really 20x Cheaper?
  32. Thunder Compute's $13M GPU Virtualization Bet: Enterprise Buyer's Guide
  33. OpenViking Review 2026: Is Filesystem Memory Production-Ready?
  34. LLM-as-a-Verifier Explained: Architecture, Costs, and Production Fit
  35. AirLLM on 4 GB VRAM: How Layer-Wise Inference Really Works
  36. TrueForge Review: Is the Open-Source Agent Harness Production-Ready?
  37. Agent-Readable Websites: llms.txt, Markdown Mirrors and What Breaks
  38. Qwen3.8-27B: Self-Hosted Computer-Use Agents Without Exporting Screenshots
  39. Localized URL Slugs vs hreflang: Why We Keep One English Slug
  40. Can an AI Agent Use Your Product, or Only Read About It?
  41. Graft Review 2026: Do Agent Repo Maps Belong in Git?
  42. How Coding Agents Keep Token Bills in Check with Output Compression
  43. Smarter Token Usage with Your AI Coding Agent
  44. DeepSeek Harness Review: Is the Plugin Stack Production-Ready?
  45. Netflix's vLLM and Triton Stack: 7 Production Lessons
  46. OpenSandbox Review: Is Self-Hosting Worth It?
  47. Transformers.js Browser AI: When Local Inference Belongs in Your Product
  48. Cloudflare Kitesurf Review: Cost, Limits and Production Fit
  49. GitHub Spec Kit Review: Is It Worth the Process?
  50. Internal AI Agent Marketplace: A 2026 Enterprise Build Guide
  51. How to Self-Host LiteLLM in Production: 2026 Guide
  52. AI-Ready Company Wiki: Architecture and Build Guide
  53. Is Linux the Best OS for AI Agents? A 2026 Infrastructure Guide
  54. Does Claude Watermark Text? The 2026 API Answer
  55. OpenKB Review: Knowledge Compiler vs RAG
  56. Unsloth Desktop Review: A Private Local AI Workstation?
  57. NeMo Switchyard 0.2: Agent Model Routing Without Training?
  58. Firecrawl AnyDoc Review: 14 Formats to Markdown
  59. MCP Cloud vs Manufact Cloud: MCP Hosting Guide
  60. Muse Glimmer 30B: Is Meta's Local Agent Model Production-Ready?
  61. How to Make AI Writing Sound Human with Agent Skills
  62. NVIDIA NOOA Review: Are Object-Oriented Agents Production-Ready?
  63. OmniRoute AI Routing: Setup and Production Checklist
  64. Strix AI Pentesting: 30-Day Pilot and Buying Guide for 2026
  65. Hyperagent Review: Cloud AI Agents Without a Server
  66. Gemini Robotics 2: Whole-Body Control and the Pilot Decision
  67. Hark Handoff Review: The Agent That Actually Clicks
  68. Meta Muse Code Pricing: Is the Contributor Tier Safe for Client Code?
  69. pdf-inspector Review: Route PDFs Before OCR
  70. PII Redaction Before LLM Prompts: A Practical Pipeline
  71. Agent Reach Review: Costs, Security and Real Limits
  72. Cloudflare Wallets for AI Agents: What Is Live?
  73. Local Multimodal AI Coding Assistant: Voice, OCR and Privacy
  74. DeepSeek V4 Flash 0731 on One AI PC: What Actually Works?
  75. YC QM Agent Review: Is Quartermaster Ready for Work?
  76. jcode vs Claude Code: Is the Rust Harness Worth Switching To?
  77. Lightpanda Browser for AI Agents: Production Guide
  78. Graph Engineering for AI Agents: When Does a Knowledge Graph Pay Off?
  79. Multi-Model AI Coding Agent Stack: A Team Buying Guide
  80. Is MCP Stateless Now? Your Server Migration Checklist
  81. MCP Is Not a Security Boundary: Protect Agent Data
  82. Fine-Tune Gemma 4 Free with Unsloth and Colab
  83. Can AI Agents Talk to Each Other? A Band Setup Guide
  84. Taalas HC1 Review: Is a Hardwired LLM ASIC Worth It?
  85. LeanCTX Technical Field Report: 64.1% Less Context
  86. llmfit Guide: Which Local LLM Fits Your Hardware?
  87. CLIProxyAPI: Run GPT-5.6 Sol Inside Claude Code
  88. Miso TTS Self-Hosted vs API: Cost, Latency and VRAM Reality
  89. Enterprise MCP Authorization Architecture
  90. Compress RAG Vectors to 2 Bits: Is Data-Oblivious Quantization Production-Ready?
  91. LMCache Reported 13.7x Lower Mean TTFT with a Shared KV Cache
  92. Lossless LLM Weight Compression vs 8-bit GGUF: What Is Ready for Production?
  93. OpenAI Eval Sandbox Escape: 12 Controls Before You Test a Cyber Agent
  94. Cisco Antares Review: 1B Local AI for Vulnerability Localization
  95. Meterless Review 2026: Is This AI Agent Context Layer Ready?
  96. NVIDIA Nemotron 3.5 ASR: Is Free Self-Hosted STT Ready for Voice Agents?
  97. Claude Code vs OpenCode: Which Costs Less for a Team in 2026?
  98. Kimi K3 for EU Companies: API Cost, Data Risk, and a Pilot Plan
  99. Mesh LLM Review: Can One Large LLM Run Across Multiple Computers?
  100. Graphify Review 2026: Is a Codebase Knowledge Graph Worth It?
  101. Bonsai 27B Review: Can a 27B LLM Really Run on a Phone?
  102. Soofi S: Is Germany's Sovereign LLM Ready for Business?
  103. AI Pilot Kill-or-Scale Scorecard: 12 Metrics to Check After 30 Days
  104. T3MP3ST Review 2026: Can It Replace a Penetration Test?
  105. Colibri Runs GLM-5.2 on Consumer Hardware. Here Is the Catch.
  106. Open Knowledge Format (OKF): The Enterprise Guide
  107. AI Agent Cost per Action: Why Agentic Workflows Blow Up Token Bills
  108. When Local Models Beat APIs: A Break-Even Calculator for EU Companies
  109. LLM Cost Calculator 2026: Cost per Task, Not Cost per Token
  110. Pxpipe Review: Can Images Cut Claude Code Token Costs 60%?
  111. Cheaper Per Token Can Still Cost More Per Task
  112. How to Use Claude Fable 5.1 in Claude Code
  113. Anthropic Says No Open-Weights Ban. The Cost War Is Still Real.
  114. What an Internal AI Assistant Actually Costs in the DACH Region (2026)
  115. LLM Gateways Compared 2026: LiteLLM vs OpenRouter vs Portkey vs RouteLLM
  116. Self-Hosting LLMs in the EU: When Open Weights Actually Pay Off
  117. Best Open-Weight LLMs 2026: DeepSeek vs Qwen vs Kimi vs GLM vs Llama
  118. The Bottleneck Was Never Intelligence. It Was Context.
  119. ChatGPT Enterprise vs Copilot vs Custom RAG
  120. MCP vs RAG vs Agent Skills vs Custom GPTs
  121. How to Cut LLM Token Costs in 2026
  122. AI Agent Pilot in 30/60/90 Days
  123. Permissions-First RAG over SharePoint, Confluence, Drive
  124. RAG Production-Readiness Checklist for EU Companies
  125. When Is an LLM Eval Worth Building? Cost, ROI, and Trusting the Judge
  126. Why AI Agent Projects Get Cancelled
  127. RAG vs Fine-Tuning vs Long-Context 2026
  128. LLM API Costs 2026. Architecture Shift