Topic

AI and agents

Engineering, operating and evaluating AI systems, coding agents and model infrastructure.

Blog posts
147
Focused clusters
02

Focused clusters

Start with the cornerstone

Inbox, without the noise

Get the next AI and agents field note

One concise email when we publish. No tracking pixels, and no inbox filler.

Latest in this collection

Layered context cards illustrating an editable agent working set AI & Agents

Context Language Models vs Compaction: What to Pilot

Should your agent edit its own context? Compare CLMs with compaction, inspect the real harness and license, and test what happens when essential facts disappear.

Wavect editorial illustration for LiteLLM Lens agent trace analysis AI & Agents

LiteLLM Lens: Agent Trace Analysis with SQL and APIs

Move from individual agent traces to evidence-linked investigations. A practical guide to LiteLLM Lens, ClickHouse SQL, safe API access and regression testing.

SmythOS Studio self-hosting guide: visual workflows, local infrastructure and runtime boundaries AI & Agents

SmythOS Studio Self-Hosting: Docker Setup, Costs and Limits

An MIT-licensed visual agent builder on your infrastructure. A source-based guide to Docker setup, real costs and the gap between a local demo and a public service.

DeerFlow agent harness: connected context, sandbox and workflow layers AI & Agents

DeerFlow 2.0: Docker Setup, Sandboxes and Memory

A practical DeerFlow guide beyond the launch claims: configure Docker, understand shared sandboxes, troubleshoot memory and define production acceptance checks.

Separate state and action paths meeting at a contrastive ranking step AI & Agents

CLM-8B Self-Hosting: vLLM, Action Cache and Verifiers

A source-checked deployment guide to CLM-8B: the full encoder behind the small heads, reusable candidates, private APIs and verifier evaluation limits.

Abstract option-score bars illustrating decision calibration, not measured benchmark values AI & Agents

AnyJev: LLM Calibration and Option-Order Bias

A deeper look at Nokia’s AnyJev: why stable answers are not calibrated probabilities, what version 0.2.0 changes and how to evaluate a real router.

Complete article directory

  1. Context Language Models vs Compaction: What to Pilot
  2. LiteLLM Lens: Agent Trace Analysis with SQL and APIs
  3. SmythOS Studio Self-Hosting: Docker Setup, Costs and Limits
  4. DeerFlow 2.0: Docker Setup, Sandboxes and Memory
  5. CLM-8B Self-Hosting: vLLM, Action Cache and Verifiers
  6. AnyJev: LLM Calibration and Option-Order Bias
  7. LiteAgents SDK: Per-Turn Routing, Setup and Migration
  8. mcp-memory-service: Shared Memory for Claude Code and Cursor
  9. Claude Opus 5.5: Best Uses, Prompts and Effort Settings
  10. Laya and Jev for Business: 6 Practical Workflows
  11. Laya vs Jev: What the Benchmarks Mean for AI Startups
  12. Google AX Agent Executor: Budgets and Self-Hosting
  13. WikiSkill: How Agents Evolve SKILL.md from Experience
  14. Tencent Octop Review: What Builders Actually Get for Free
  15. Wigolo Review: Local Web Intelligence for AI Agents
  16. Jev AI Review: Decision Models for Agent Workflows
  17. Supermemory: AI Agent Memory, RAG and Local Setup
  18. AI Agent Design Patterns: Start Simple, Verify Actions
  19. Claude Mods: Setup, Function Hooks and Security
  20. Voicebox: Local Voice Cloning, Dictation and MCP Setup
  21. Valyu’s 0.6B Multi-Agent Router: Results, Limits and When to Train One
  22. OpenAI Agents API Review: Migration, Costs and Data Controls
  23. Spotify shunt Review: Setup, Savings and Limits
  24. OpenBot Review: Self-Hosted AI Coworkers, Costs & Controls
  25. SwarmLLM Review 2026: Browser P2P LLM Inference Across Phones and Laptops
  26. Ramp Inspect Architecture 2026: Background Coding Agents at Scale
  27. Model Hardware Standard: Enterprise Guide to Physical AI
  28. Fonio AI Review 2026: Pricing, API, GDPR & Build vs Buy
  29. Mosaic (YC S26) Review: Shared Memory for Team AI Agents
  30. Ripwire Review 2026: AI Repo Context Without Embeddings?
  31. Atomic Multi-File Edits for AI Coding Agents: The Semaprax Lesson
  32. Phonely Alma Review: Is the Voice LLM Ready for Production?
  33. Utopia Review: Temporal Knowledge Graph for Enterprise
  34. AI Agent Knowledge Transfer: Use Frontier Models Once, Then Scale Cheaper
  35. claude-rotate: One Proxy for Multiple Claude Max Accounts
  36. NVIDIA PAIR Review: Local AI Routing and the AMD Gap
  37. Feynman Review: Is This Open-Source AI Research Agent Ready for Teams?
  38. Obscura Browser Review: Claims, Limits and Production Fit
  39. Claude Code Design System: 4 Parts for On-Brand UI
  40. Tencent Hy4 Preview Review: Is the 1M-Context Coding Model Worth a Pilot?
  41. PageLM Review: Self-Hosting and Commercial Use
  42. ChatGPT Can Now Log In Without Seeing Your Password
  43. M6 Mac mini vs M5 Mac Studio for Local AI: Buyer's Guide
  44. AI Agent Harness, Explained: The Reliability Layer Around an LLM
  45. FreeToken AI Review: Run Frontier MoE Models on Consumer GPUs?
  46. Darkbloom AI Review: Private Inference on Idle Macs
  47. Ox Alpha Revealed as GLM-5.3-Flash: What Changed
  48. Why Agent Edits Need Semantic Identity: Building SEMAPRAX in Rust
  49. LangChain Deep Agents Review: Is the Agent Harness Ready for Production?
  50. Pika Audio Models API Pricing: Is SFX Really 20x Cheaper?
  51. Thunder Compute's $13M GPU Virtualization Bet: Enterprise Buyer's Guide
  52. OpenViking Review: Filesystem Memory for AI Agents
  53. LLM-as-a-Verifier Explained: Architecture, Costs, and Production Fit
  54. AirLLM on 4 GB VRAM: How Layer-Wise Inference Really Works
  55. TrueForge Review: Is the Open-Source Agent Harness Production-Ready?
  56. Agent-Readable Websites: llms.txt, Markdown Mirrors and What Breaks
  57. Qwen3.8-27B: Self-Hosted Computer-Use Agents Without Exporting Screenshots
  58. Localized URL Slugs vs hreflang: Why We Keep One English Slug
  59. Can an AI Agent Use Your Product, or Only Read About It?
  60. Graft Review 2026: Do Agent Repo Maps Belong in Git?
  61. How Coding Agents Keep Token Bills in Check with Output Compression
  62. Smarter Token Usage with Your AI Coding Agent
  63. DeepSeek Harness Review: Is the Plugin Stack Production-Ready?
  64. Netflix's vLLM and Triton Stack: 7 Production Lessons
  65. OpenSandbox Review: Is Self-Hosting Worth It?
  66. Transformers.js Browser AI: When Local Inference Belongs in Your Product
  67. Cloudflare Kitesurf Review: Cost, Limits and Production Fit
  68. GitHub Spec Kit Review: Is It Worth the Process?
  69. Internal AI Agent Marketplace: A 2026 Enterprise Build Guide
  70. How to Self-Host LiteLLM in Production: 2026 Guide
  71. AI-Ready Company Wiki: Architecture and Build Guide
  72. Is Linux the Best OS for AI Agents? A 2026 Infrastructure Guide
  73. Does Claude Watermark Text? The 2026 API Answer
  74. OpenKB Review: Knowledge Compiler vs RAG
  75. Unsloth Desktop Review: A Private Local AI Workstation?
  76. NeMo Switchyard 0.2: Agent Model Routing Without Training?
  77. Firecrawl AnyDoc Review: 14 Formats to Markdown
  78. MCP Cloud vs Manufact Cloud: MCP Hosting Guide
  79. Muse Glimmer 30B: Is Meta's Local Agent Model Production-Ready?
  80. How to Make AI Writing Sound Human with Agent Skills
  81. NVIDIA NOOA Review: Are Object-Oriented Agents Production-Ready?
  82. OmniRoute AI Routing: Setup and Production Checklist
  83. Strix AI Pentesting: 30-Day Pilot and Buying Guide for 2026
  84. Hyperagent Review: Cloud AI Agents Without a Server
  85. Gemini Robotics 2: Whole-Body Control and the Pilot Decision
  86. Hark Handoff Review: The Agent That Actually Clicks
  87. Meta Muse Code Pricing: Is the Contributor Tier Safe for Client Code?
  88. pdf-inspector Review: Route PDFs Before OCR
  89. PII Redaction Before LLM Prompts: A Practical Pipeline
  90. Agent Reach Review: Costs, Security and Real Limits
  91. Cloudflare Wallets for AI Agents: What Is Live?
  92. Local Multimodal AI Coding Assistant: Voice, OCR and Privacy
  93. DeepSeek V4 Flash 0731 on One AI PC: What Actually Works?
  94. YC QM Agent Review: Is Quartermaster Ready for Work?
  95. jcode vs Claude Code: Is the Rust Harness Worth Switching To?
  96. Lightpanda Browser for AI Agents: Production Guide
  97. Graph Engineering for AI Agents: When Does a Knowledge Graph Pay Off?
  98. Multi-Model AI Coding Agent Stack: A Team Buying Guide
  99. Is MCP Stateless Now? Your Server Migration Checklist
  100. MCP Is Not a Security Boundary: Protect Agent Data
  101. Fine-Tune Gemma 4 Free with Unsloth and Colab
  102. Can AI Agents Talk to Each Other? A Band Setup Guide
  103. Taalas HC1 Review: Is a Hardwired LLM ASIC Worth It?
  104. LeanCTX Technical Field Report: 64.1% Less Context
  105. llmfit Guide: Which Local LLM Fits Your Hardware?
  106. CLIProxyAPI: Run GPT-5.6 Sol Inside Claude Code
  107. Miso TTS Self-Hosted vs API: Cost, Latency and VRAM Reality
  108. Enterprise MCP Authorization Architecture
  109. Compress RAG Vectors to 2 Bits: Is Data-Oblivious Quantization Production-Ready?
  110. LMCache Reported 13.7x Lower Mean TTFT with a Shared KV Cache
  111. Lossless LLM Weight Compression vs 8-bit GGUF: What Is Ready for Production?
  112. OpenAI Eval Sandbox Escape: 12 Controls Before You Test a Cyber Agent
  113. Cisco Antares Review: 1B Local AI for Vulnerability Localization
  114. Meterless Review 2026: Is This AI Agent Context Layer Ready?
  115. NVIDIA Nemotron 3.5 ASR: Is Free Self-Hosted STT Ready for Voice Agents?
  116. Claude Code vs OpenCode: Which Costs Less for a Team in 2026?
  117. Kimi K3 for EU Companies: API Cost, Data Risk, and a Pilot Plan
  118. Mesh LLM Review: Can One Large LLM Run Across Multiple Computers?
  119. Graphify Review 2026: Is a Codebase Knowledge Graph Worth It?
  120. Bonsai 27B Review: Can a 27B LLM Really Run on a Phone?
  121. Soofi S: Is Germany's Sovereign LLM Ready for Business?
  122. AI Pilot Kill-or-Scale Scorecard: 12 Metrics to Check After 30 Days
  123. T3MP3ST Review 2026: Can It Replace a Penetration Test?
  124. Colibri Runs GLM-5.2 on Consumer Hardware. Here Is the Catch.
  125. Open Knowledge Format (OKF): The Enterprise Guide
  126. AI Agent Cost per Action: Why Agentic Workflows Blow Up Token Bills
  127. When Local Models Beat APIs: A Break-Even Calculator for EU Companies
  128. LLM Cost Calculator 2026: Cost per Task, Not Cost per Token
  129. Pxpipe Review: Can Images Cut Claude Code Token Costs 60%?
  130. Cheaper Per Token Can Still Cost More Per Task
  131. How to Use Claude Fable 5.1 in Claude Code
  132. Anthropic Says No Open-Weights Ban. The Cost War Is Still Real.
  133. What an Internal AI Assistant Actually Costs in the DACH Region (2026)
  134. LLM Gateways Compared 2026: LiteLLM vs OpenRouter vs Portkey vs RouteLLM
  135. Self-Hosting LLMs in the EU: When Open Weights Actually Pay Off
  136. Best Open-Weight LLMs 2026: DeepSeek vs Qwen vs Kimi vs GLM vs Llama
  137. The Bottleneck Was Never Intelligence. It Was Context.
  138. ChatGPT Enterprise vs Copilot vs Custom RAG
  139. MCP vs RAG vs Agent Skills vs Custom GPTs
  140. How to Cut LLM Token Costs in 2026
  141. AI Agent Pilot in 30/60/90 Days
  142. Permissions-First RAG over SharePoint, Confluence, Drive
  143. RAG Production-Readiness Checklist for EU Companies
  144. When Is an LLM Eval Worth Building? Cost, ROI, and Trusting the Judge
  145. Why AI Agent Projects Get Cancelled
  146. RAG vs Fine-Tuning vs Long-Context 2026
  147. LLM API Costs 2026. Architecture Shift