Models and infrastructure
Model selection, inference economics, local deployment, compression and serving architecture.
- Blog posts
- 67
- Topic
- AI and agents
Start with the cornerstone
Get the next AI and agents field note
One concise email when we publish. No tracking pixels, and no inbox filler.
Latest in this collection
Context Language Models vs Compaction: What to Pilot
Should your agent edit its own context? Compare CLMs with compaction, inspect the real harness and license, and test what happens when essential facts disappear.
LiteLLM Lens: Agent Trace Analysis with SQL and APIs
Move from individual agent traces to evidence-linked investigations. A practical guide to LiteLLM Lens, ClickHouse SQL, safe API access and regression testing.
SmythOS Studio Self-Hosting: Docker Setup, Costs and Limits
An MIT-licensed visual agent builder on your infrastructure. A source-based guide to Docker setup, real costs and the gap between a local demo and a public service.
DeerFlow 2.0: Docker Setup, Sandboxes and Memory
A practical DeerFlow guide beyond the launch claims: configure Docker, understand shared sandboxes, troubleshoot memory and define production acceptance checks.
CLM-8B Self-Hosting: vLLM, Action Cache and Verifiers
A source-checked deployment guide to CLM-8B: the full encoder behind the small heads, reusable candidates, private APIs and verifier evaluation limits.
AnyJev: LLM Calibration and Option-Order Bias
A deeper look at Nokia’s AnyJev: why stable answers are not calibrated probabilities, what version 0.2.0 changes and how to evaluate a real router.
Complete article directory
- Context Language Models vs Compaction: What to Pilot
- LiteLLM Lens: Agent Trace Analysis with SQL and APIs
- SmythOS Studio Self-Hosting: Docker Setup, Costs and Limits
- DeerFlow 2.0: Docker Setup, Sandboxes and Memory
- CLM-8B Self-Hosting: vLLM, Action Cache and Verifiers
- AnyJev: LLM Calibration and Option-Order Bias
- LiteAgents SDK: Per-Turn Routing, Setup and Migration
- mcp-memory-service: Shared Memory for Claude Code and Cursor
- Claude Opus 5.5: Best Uses, Prompts and Effort Settings
- Laya and Jev for Business: 6 Practical Workflows
- Laya vs Jev: What the Benchmarks Mean for AI Startups
- Jev AI Review: Decision Models for Agent Workflows
- SwarmLLM Review 2026: Browser P2P LLM Inference Across Phones and Laptops
- Phonely Alma Review: Is the Voice LLM Ready for Production?
- Utopia Review: Temporal Knowledge Graph for Enterprise
- NVIDIA PAIR Review: Local AI Routing and the AMD Gap
- Tencent Hy4 Preview Review: Is the 1M-Context Coding Model Worth a Pilot?
- PageLM Review: Self-Hosting and Commercial Use
- M6 Mac mini vs M5 Mac Studio for Local AI: Buyer's Guide
- FreeToken AI Review: Run Frontier MoE Models on Consumer GPUs?
- Darkbloom AI Review: Private Inference on Idle Macs
- Ox Alpha Revealed as GLM-5.3-Flash: What Changed
- Pika Audio Models API Pricing: Is SFX Really 20x Cheaper?
- Thunder Compute's $13M GPU Virtualization Bet: Enterprise Buyer's Guide
- AirLLM on 4 GB VRAM: How Layer-Wise Inference Really Works
- Qwen3.8-27B: Self-Hosted Computer-Use Agents Without Exporting Screenshots
- Netflix's vLLM and Triton Stack: 7 Production Lessons
- Transformers.js Browser AI: When Local Inference Belongs in Your Product
- How to Self-Host LiteLLM in Production: 2026 Guide
- AI-Ready Company Wiki: Architecture and Build Guide
- Does Claude Watermark Text? The 2026 API Answer
- OpenKB Review: Knowledge Compiler vs RAG
- Unsloth Desktop Review: A Private Local AI Workstation?
- NeMo Switchyard 0.2: Agent Model Routing Without Training?
- Firecrawl AnyDoc Review: 14 Formats to Markdown
- Muse Glimmer 30B: Is Meta's Local Agent Model Production-Ready?
- OmniRoute AI Routing: Setup and Production Checklist
- Gemini Robotics 2: Whole-Body Control and the Pilot Decision
- pdf-inspector Review: Route PDFs Before OCR
- Local Multimodal AI Coding Assistant: Voice, OCR and Privacy
- DeepSeek V4 Flash 0731 on One AI PC: What Actually Works?
- Fine-Tune Gemma 4 Free with Unsloth and Colab
- Taalas HC1 Review: Is a Hardwired LLM ASIC Worth It?
- llmfit Guide: Which Local LLM Fits Your Hardware?
- Miso TTS Self-Hosted vs API: Cost, Latency and VRAM Reality
- Compress RAG Vectors to 2 Bits: Is Data-Oblivious Quantization Production-Ready?
- LMCache Reported 13.7x Lower Mean TTFT with a Shared KV Cache
- Lossless LLM Weight Compression vs 8-bit GGUF: What Is Ready for Production?
- Cisco Antares Review: 1B Local AI for Vulnerability Localization
- NVIDIA Nemotron 3.5 ASR: Is Free Self-Hosted STT Ready for Voice Agents?
- Kimi K3 for EU Companies: API Cost, Data Risk, and a Pilot Plan
- Mesh LLM Review: Can One Large LLM Run Across Multiple Computers?
- Bonsai 27B Review: Can a 27B LLM Really Run on a Phone?
- Soofi S: Is Germany's Sovereign LLM Ready for Business?
- Colibri Runs GLM-5.2 on Consumer Hardware. Here Is the Catch.
- When Local Models Beat APIs: A Break-Even Calculator for EU Companies
- LLM Cost Calculator 2026: Cost per Task, Not Cost per Token
- Pxpipe Review: Can Images Cut Claude Code Token Costs 60%?
- Cheaper Per Token Can Still Cost More Per Task
- Anthropic Says No Open-Weights Ban. The Cost War Is Still Real.
- LLM Gateways Compared 2026: LiteLLM vs OpenRouter vs Portkey vs RouteLLM
- Self-Hosting LLMs in the EU: When Open Weights Actually Pay Off
- Best Open-Weight LLMs 2026: DeepSeek vs Qwen vs Kimi vs GLM vs Llama
- How to Cut LLM Token Costs in 2026
- When Is an LLM Eval Worth Building? Cost, ROI, and Trusting the Judge
- RAG vs Fine-Tuning vs Long-Context 2026
- LLM API Costs 2026. Architecture Shift