AI and agents
Engineering, operating and evaluating AI systems, coding agents and model infrastructure.
- Blog posts
- 127
- Focused clusters
- 02
Focused clusters
Start with the cornerstone
- Graph Engineering for AI Agents: When Does a Knowledge Graph Pay Off?A buyer-focused guide to execution graphs, experiment DAGs and knowledge graphs, including when to skip them, what provenance requires and how to run a measured pilot.
- Self-Hosting LLMs in the EU: When Open Weights Actually Pay OffThe GPU is the cheap part. Here is the real cost of self-hosting open weights, a defensible break-even formula, the GDPR trade-offs, and the vLLM production stack.
Get the next AI and agents field note
One concise email when we publish. No tracking pixels, and no inbox filler.
Latest in this collection
AI & AgentsValyu’s 0.6B Multi-Agent Router: Results, Limits and When to Train One
A trained Qwen3-0.6B beat prompted Haiku and Nova Lite in one retrieval-routing study. Understand the metric, training evidence and production decision before adopting it.
AI & AgentsOpenAI Agents API Review: Migration, Costs and Data Controls
What the managed Codex harness replaces, what stays your responsibility, and how to test the launch numbers before migrating a real workflow.
AI & AgentsSpotify shunt Review: Setup, Savings and Limits
What Spotify's 90% claim actually measures, what Portal setup requires, and how to test file-read delegation without moving the cost into rework.
AI & AgentsOpenBot Review: Self-Hosted AI Coworkers, Costs & Controls
A source-reviewed guide to CopilotKit OpenBot: where its action gateway helps, which defaults need changing, and when self-hosting beats buying a managed assistant.
SwarmLLM Review 2026: Browser P2P LLM Inference Across Phones and Laptops
A buyer-focused review of SwarmLLM: browser-native 27B inference across phones and laptops, WebGPU/WebRTC architecture, benchmark limits, peer trust, and when to pilot it.
Ramp Inspect Architecture 2026: Background Coding Agents at Scale
How Ramp Inspect uses full-stack Modal sandboxes, filesystem snapshots, coordination and verification, plus the controls a regulated engineering team needs before scaling background coding agents.
Complete article directory
- Valyu’s 0.6B Multi-Agent Router: Results, Limits and When to Train One
- OpenAI Agents API Review: Migration, Costs and Data Controls
- Spotify shunt Review: Setup, Savings and Limits
- OpenBot Review: Self-Hosted AI Coworkers, Costs & Controls
- SwarmLLM Review 2026: Browser P2P LLM Inference Across Phones and Laptops
- Ramp Inspect Architecture 2026: Background Coding Agents at Scale
- Model Hardware Standard: Enterprise Guide to Physical AI
- Fonio AI Review 2026: Pricing, API, GDPR & Build vs Buy
- Mosaic (YC S26) Review: Shared Memory for Team AI Agents
- Ripwire Review 2026: AI Repo Context Without Embeddings?
- Atomic Multi-File Edits for AI Coding Agents: The Semaprax Lesson
- Phonely Alma Review: Is the Voice LLM Ready for Production?
- Utopia Review: Temporal Knowledge Graph for Enterprise
- AI Agent Knowledge Transfer: Use Frontier Models Once, Then Scale Cheaper
- claude-rotate: One Proxy for Multiple Claude Max Accounts
- NVIDIA PAIR Review: Local AI Routing and the AMD Gap
- Feynman Review: Is This Open-Source AI Research Agent Ready for Teams?
- Obscura Browser Review: Claims, Limits and Production Fit
- Claude Code Design System: 4 Parts for On-Brand UI
- Tencent Hy4 Preview Review: Is the 1M-Context Coding Model Worth a Pilot?
- PageLM Review: Self-Hosting and Commercial Use
- ChatGPT Can Now Log In Without Seeing Your Password
- M6 Mac mini vs M5 Mac Studio for Local AI: Buyer's Guide
- AI Agent Harness, Explained: The Reliability Layer Around an LLM
- FreeToken AI Review: Run Frontier MoE Models on Consumer GPUs?
- Darkbloom AI Review: Private Inference on Idle Macs
- Ox Alpha Revealed as GLM-5.3-Flash: What Changed
- Why Agent Edits Need Semantic Identity: Building SEMAPRAX in Rust
- LangChain Deep Agents Review: Is the Agent Harness Ready for Production?
- Pika Audio Models API Pricing: Is SFX Really 20x Cheaper?
- Thunder Compute's $13M GPU Virtualization Bet: Enterprise Buyer's Guide
- OpenViking Review 2026: Is Filesystem Memory Production-Ready?
- LLM-as-a-Verifier Explained: Architecture, Costs, and Production Fit
- AirLLM on 4 GB VRAM: How Layer-Wise Inference Really Works
- TrueForge Review: Is the Open-Source Agent Harness Production-Ready?
- Agent-Readable Websites: llms.txt, Markdown Mirrors and What Breaks
- Qwen3.8-27B: Self-Hosted Computer-Use Agents Without Exporting Screenshots
- Localized URL Slugs vs hreflang: Why We Keep One English Slug
- Can an AI Agent Use Your Product, or Only Read About It?
- Graft Review 2026: Do Agent Repo Maps Belong in Git?
- How Coding Agents Keep Token Bills in Check with Output Compression
- Smarter Token Usage with Your AI Coding Agent
- DeepSeek Harness Review: Is the Plugin Stack Production-Ready?
- Netflix's vLLM and Triton Stack: 7 Production Lessons
- OpenSandbox Review: Is Self-Hosting Worth It?
- Transformers.js Browser AI: When Local Inference Belongs in Your Product
- Cloudflare Kitesurf Review: Cost, Limits and Production Fit
- GitHub Spec Kit Review: Is It Worth the Process?
- Internal AI Agent Marketplace: A 2026 Enterprise Build Guide
- How to Self-Host LiteLLM in Production: 2026 Guide
- AI-Ready Company Wiki: Architecture and Build Guide
- Is Linux the Best OS for AI Agents? A 2026 Infrastructure Guide
- Does Claude Watermark Text? The 2026 API Answer
- OpenKB Review: Knowledge Compiler vs RAG
- Unsloth Desktop Review: A Private Local AI Workstation?
- NeMo Switchyard 0.2: Agent Model Routing Without Training?
- Firecrawl AnyDoc Review: 14 Formats to Markdown
- MCP Cloud vs Manufact Cloud: MCP Hosting Guide
- Muse Glimmer 30B: Is Meta's Local Agent Model Production-Ready?
- How to Make AI Writing Sound Human with Agent Skills
- NVIDIA NOOA Review: Are Object-Oriented Agents Production-Ready?
- OmniRoute AI Routing: Setup and Production Checklist
- Strix AI Pentesting: 30-Day Pilot and Buying Guide for 2026
- Hyperagent Review: Cloud AI Agents Without a Server
- Gemini Robotics 2: Whole-Body Control and the Pilot Decision
- Hark Handoff Review: The Agent That Actually Clicks
- Meta Muse Code Pricing: Is the Contributor Tier Safe for Client Code?
- pdf-inspector Review: Route PDFs Before OCR
- PII Redaction Before LLM Prompts: A Practical Pipeline
- Agent Reach Review: Costs, Security and Real Limits
- Cloudflare Wallets for AI Agents: What Is Live?
- Local Multimodal AI Coding Assistant: Voice, OCR and Privacy
- DeepSeek V4 Flash 0731 on One AI PC: What Actually Works?
- YC QM Agent Review: Is Quartermaster Ready for Work?
- jcode vs Claude Code: Is the Rust Harness Worth Switching To?
- Lightpanda Browser for AI Agents: Production Guide
- Graph Engineering for AI Agents: When Does a Knowledge Graph Pay Off?
- Multi-Model AI Coding Agent Stack: A Team Buying Guide
- Is MCP Stateless Now? Your Server Migration Checklist
- MCP Is Not a Security Boundary: Protect Agent Data
- Fine-Tune Gemma 4 Free with Unsloth and Colab
- Can AI Agents Talk to Each Other? A Band Setup Guide
- Taalas HC1 Review: Is a Hardwired LLM ASIC Worth It?
- LeanCTX Technical Field Report: 64.1% Less Context
- llmfit Guide: Which Local LLM Fits Your Hardware?
- CLIProxyAPI: Run GPT-5.6 Sol Inside Claude Code
- Miso TTS Self-Hosted vs API: Cost, Latency and VRAM Reality
- Enterprise MCP Authorization Architecture
- Compress RAG Vectors to 2 Bits: Is Data-Oblivious Quantization Production-Ready?
- LMCache Reported 13.7x Lower Mean TTFT with a Shared KV Cache
- Lossless LLM Weight Compression vs 8-bit GGUF: What Is Ready for Production?
- OpenAI Eval Sandbox Escape: 12 Controls Before You Test a Cyber Agent
- Cisco Antares Review: 1B Local AI for Vulnerability Localization
- Meterless Review 2026: Is This AI Agent Context Layer Ready?
- NVIDIA Nemotron 3.5 ASR: Is Free Self-Hosted STT Ready for Voice Agents?
- Claude Code vs OpenCode: Which Costs Less for a Team in 2026?
- Kimi K3 for EU Companies: API Cost, Data Risk, and a Pilot Plan
- Mesh LLM Review: Can One Large LLM Run Across Multiple Computers?
- Graphify Review 2026: Is a Codebase Knowledge Graph Worth It?
- Bonsai 27B Review: Can a 27B LLM Really Run on a Phone?
- Soofi S: Is Germany's Sovereign LLM Ready for Business?
- AI Pilot Kill-or-Scale Scorecard: 12 Metrics to Check After 30 Days
- T3MP3ST Review 2026: Can It Replace a Penetration Test?
- Colibri Runs GLM-5.2 on Consumer Hardware. Here Is the Catch.
- Open Knowledge Format (OKF): The Enterprise Guide
- AI Agent Cost per Action: Why Agentic Workflows Blow Up Token Bills
- When Local Models Beat APIs: A Break-Even Calculator for EU Companies
- LLM Cost Calculator 2026: Cost per Task, Not Cost per Token
- Pxpipe Review: Can Images Cut Claude Code Token Costs 60%?
- Cheaper Per Token Can Still Cost More Per Task
- How to Use Claude Fable 5.1 in Claude Code
- Anthropic Says No Open-Weights Ban. The Cost War Is Still Real.
- What an Internal AI Assistant Actually Costs in the DACH Region (2026)
- LLM Gateways Compared 2026: LiteLLM vs OpenRouter vs Portkey vs RouteLLM
- Self-Hosting LLMs in the EU: When Open Weights Actually Pay Off
- Best Open-Weight LLMs 2026: DeepSeek vs Qwen vs Kimi vs GLM vs Llama
- The Bottleneck Was Never Intelligence. It Was Context.
- ChatGPT Enterprise vs Copilot vs Custom RAG
- MCP vs RAG vs Agent Skills vs Custom GPTs
- How to Cut LLM Token Costs in 2026
- AI Agent Pilot in 30/60/90 Days
- Permissions-First RAG over SharePoint, Confluence, Drive
- RAG Production-Readiness Checklist for EU Companies
- When Is an LLM Eval Worth Building? Cost, ROI, and Trusting the Judge
- Why AI Agent Projects Get Cancelled
- RAG vs Fine-Tuning vs Long-Context 2026
- LLM API Costs 2026. Architecture Shift