AI and agents
Engineering, operating and evaluating AI systems, coding agents and model infrastructure.
- Blog posts
- 101
- Focused clusters
- 02
Focused clusters
Start with the cornerstone
- Graph Engineering for AI Agents: When Does a Knowledge Graph Pay Off?A buyer-focused guide to execution graphs, experiment DAGs and knowledge graphs, including when to skip them, what provenance requires and how to run a measured pilot.
- Self-Hosting LLMs in the EU: When Open Weights Actually Pay OffThe GPU is the cheap part. Here is the real cost of self-hosting open weights, the tokens/day break-even vs hosted APIs, when data residency forces your hand, and the vLLM production stack.
Latest in this collection
Ox Alpha Free AI Model: Setup, Privacy and Buyer Guide
Use the 1M-context stealth coding model through OpenCode or OpenRouter. Compare IDs, privacy terms, hidden costs and a safe team evaluation plan.
Why Agent Edits Need Semantic Identity: Building SEMAPRAX in Rust
A compiler engineering walkthrough of persistent IDs, checked HIR, deterministic graphs and replayable patches in the experimental SEMAPRAX language.
LangChain Deep Agents Review: Is the Agent Harness Ready for Production?
A buyer-focused Deep Agents review covering architecture, v0.7 changes, security boundaries, operating costs and a measurable production pilot.
Pika Audio Models API Pricing: Is SFX Really 20x Cheaper?
Pika lists SFX at $0.0002 per second, plus three more audio models. See the real break-even, API limits, quality gates and production fit before you switch.
Thunder Compute's $13M GPU Virtualization Bet: Enterprise Buyer's Guide
Thunder Compute says GPU pooling can recover stranded capacity. Compare network virtualization with MIG, vGPU and passthrough, then scope a measurable enterprise pilot.
OpenViking Review 2026: Is Filesystem Memory Production-Ready?
Buyer review of OpenViking for agent memory and RAG: architecture, AGPL risk, security controls, operating cost and a measurable two-week pilot.
Complete article directory
- Ox Alpha Free AI Model: Setup, Privacy and Buyer Guide
- Why Agent Edits Need Semantic Identity: Building SEMAPRAX in Rust
- LangChain Deep Agents Review: Is the Agent Harness Ready for Production?
- Pika Audio Models API Pricing: Is SFX Really 20x Cheaper?
- Thunder Compute's $13M GPU Virtualization Bet: Enterprise Buyer's Guide
- OpenViking Review 2026: Is Filesystem Memory Production-Ready?
- LLM-as-a-Verifier Explained: Architecture, Costs, and Production Fit
- AirLLM on 4 GB VRAM: How Layer-Wise Inference Really Works
- TrueForge Review: Is the Open-Source Agent Harness Production-Ready?
- Agent-Readable Websites: llms.txt, Markdown Mirrors and What Breaks
- Qwen3.8-27B: Self-Hosted Computer-Use Agents Without Exporting Screenshots
- Localized URLs Break hreflang: Keep One English Slug
- Can an AI Agent Use Your Product, or Only Read About It?
- Graft Review 2026: Do Agent Repo Maps Belong in Git?
- How Coding Agents Keep Token Bills in Check with Output Compression
- Smarter Token Usage with Your AI Coding Agent
- DeepSeek Harness Review: Is the Plugin Stack Production-Ready?
- Netflix's vLLM and Triton Stack: 7 Production Lessons
- OpenSandbox Review: Is Self-Hosting Worth It?
- Transformers.js Browser AI: When Local Inference Belongs in Your Product
- Cloudflare Kitesurf Review: Cost, Limits and Production Fit
- GitHub Spec Kit Review: Is It Worth the Process?
- Internal AI Agent Marketplace: A 2026 Enterprise Build Guide
- How to Self-Host LiteLLM in Production: 2026 Guide
- AI-Ready Company Wiki: Architecture and Build Guide
- Is Linux the Best OS for AI Agents? A 2026 Infrastructure Guide
- Does Claude Watermark Text? The 2026 API Answer
- OpenKB Review: Knowledge Compiler vs RAG
- Unsloth Desktop Review: A Private Local AI Workstation?
- NeMo Switchyard 0.2: Agent Model Routing Without Training?
- Firecrawl AnyDoc Review: 14 Formats to Markdown
- MCP Cloud vs Manufact Cloud: MCP Hosting Guide
- Muse Glimmer 30B: Is Meta's Local Agent Model Production-Ready?
- How to Make AI Writing Sound Human with Agent Skills
- NVIDIA NOOA Review: Are Object-Oriented Agents Production-Ready?
- OmniRoute AI Routing: Setup and Production Checklist
- Strix AI Pentesting: 30-Day Pilot and Buying Guide for 2026
- Hyperagent Review: Cloud AI Agents Without a Server
- Gemini Robotics 2: Whole-Body Control and the Pilot Decision
- Hark Handoff Review: The Agent That Actually Clicks
- Meta Muse Code Pricing: Is the Contributor Tier Safe for Client Code?
- pdf-inspector Review: Route PDFs Before OCR
- PII Redaction Before LLM Prompts: A Practical Pipeline
- Agent Reach Review: Costs, Security and Real Limits
- Cloudflare Wallets for AI Agents: What Is Live?
- Local Multimodal AI Coding Assistant: Voice, OCR and Privacy
- DeepSeek V4 Flash 0731 on One AI PC: What Actually Works?
- YC QM Agent Review: Is Quartermaster Ready for Work?
- jcode vs Claude Code: Is the Rust Harness Worth Switching To?
- Lightpanda Browser for AI Agents: Production Guide
- Graph Engineering for AI Agents: When Does a Knowledge Graph Pay Off?
- Multi-Model AI Coding Agent Stack: A Team Buying Guide
- Is MCP Stateless Now? Your Server Migration Checklist
- MCP Is Not a Security Boundary: Protect Agent Data
- Fine-Tune Gemma 4 Free with Unsloth and Colab
- Can AI Agents Talk to Each Other? A Band Setup Guide
- Taalas HC1 Review: Is a Hardwired LLM ASIC Worth It?
- LeanCTX Technical Field Report: 64.1% Less Context
- llmfit Guide: Which Local LLM Fits Your Hardware?
- CLIProxyAPI: Run GPT-5.6 Sol Inside Claude Code
- Miso TTS Self-Hosted vs API: Cost, Latency and VRAM Reality
- Enterprise MCP Authorization Architecture
- Cut RAG Vector Memory 16x: Is Data-Oblivious Quantization Ready for Production?
- Shared KV Cache Cut LLM Inference Latency 14x, With No New GPUs
- Lossless LLM Weight Compression vs 8-bit GGUF: What Is Ready for Production?
- OpenAI Eval Sandbox Escape: 12 Controls Before You Test a Cyber Agent
- Cisco Antares Review: 1B Local AI for Vulnerability Localization
- Meterless Review 2026: Is This AI Agent Context Layer Ready?
- NVIDIA Nemotron 3.5 ASR: Is Free Self-Hosted STT Ready for Voice Agents?
- Claude Code vs OpenCode: Which Costs Less for a Team in 2026?
- Kimi K3 for EU Companies: API Cost, Data Risk, and a Pilot Plan
- Mesh LLM Review: Can One Large LLM Run Across Multiple Computers?
- Graphify Review 2026: Is a Codebase Knowledge Graph Worth It?
- Bonsai 27B Review: Can a 27B LLM Really Run on a Phone?
- Soofi S: Is Germany's Sovereign LLM Ready for Business?
- AI Pilot Kill-or-Scale Scorecard: 12 Metrics to Check After 30 Days
- T3MP3ST Review 2026: Can It Replace a Penetration Test?
- Colibri Runs GLM-5.2 on Consumer Hardware. Here Is the Catch.
- Open Knowledge Format (OKF): The Enterprise Guide
- AI Agent Cost per Action: Why Agentic Workflows Blow Up Token Bills
- When Local Models Beat APIs: A Break-Even Calculator for EU Companies
- LLM Cost Calculator 2026: Cost per Task, Not Cost per Token
- Pxpipe Review: Can Images Cut Claude Code Token Costs 60%?
- Cheaper Per Token. More Expensive Per Answer.
- How to Use Claude Fable 5 in Claude Code
- Dario Declared War on Open Source. The Real War Is Over Your AI Bill.
- What an Internal AI Assistant Actually Costs in the DACH Region (2026)
- LLM Gateways Compared 2026: LiteLLM vs OpenRouter vs Portkey vs RouteLLM
- Self-Hosting LLMs in the EU: When Open Weights Actually Pay Off
- Best Open-Weight LLMs 2026: DeepSeek vs Qwen vs Kimi vs GLM vs Llama
- The Bottleneck Was Never Intelligence. It Was Context.
- ChatGPT Enterprise vs Copilot vs Custom RAG
- MCP vs RAG vs Agent Skills vs Custom GPTs
- How to Cut LLM Token Costs in 2026
- AI Agent Pilot in 30/60/90 Days
- Permissions-First RAG over SharePoint, Confluence, Drive
- RAG Production-Readiness Checklist for EU Companies
- When Is an LLM Eval Worth Building? Cost, ROI, and Trusting the Judge
- Why 40% of AI Agent Projects Die
- RAG vs Fine-Tuning vs Long-Context 2026
- LLM API Costs 2026. Architecture Shift