Cluster

Models and infrastructure

Model selection, inference economics, local deployment, compression and serving architecture.

Blog posts
67
Topic
AI and agents

Start with the cornerstone

Inbox, without the noise

Get the next AI and agents field note

One concise email when we publish. No tracking pixels, and no inbox filler.

Latest in this collection

Layered context cards illustrating an editable agent working set AI & Agents

Context Language Models vs Compaction: What to Pilot

Should your agent edit its own context? Compare CLMs with compaction, inspect the real harness and license, and test what happens when essential facts disappear.

Wavect editorial illustration for LiteLLM Lens agent trace analysis AI & Agents

LiteLLM Lens: Agent Trace Analysis with SQL and APIs

Move from individual agent traces to evidence-linked investigations. A practical guide to LiteLLM Lens, ClickHouse SQL, safe API access and regression testing.

SmythOS Studio self-hosting guide: visual workflows, local infrastructure and runtime boundaries AI & Agents

SmythOS Studio Self-Hosting: Docker Setup, Costs and Limits

An MIT-licensed visual agent builder on your infrastructure. A source-based guide to Docker setup, real costs and the gap between a local demo and a public service.

DeerFlow agent harness: connected context, sandbox and workflow layers AI & Agents

DeerFlow 2.0: Docker Setup, Sandboxes and Memory

A practical DeerFlow guide beyond the launch claims: configure Docker, understand shared sandboxes, troubleshoot memory and define production acceptance checks.

Separate state and action paths meeting at a contrastive ranking step AI & Agents

CLM-8B Self-Hosting: vLLM, Action Cache and Verifiers

A source-checked deployment guide to CLM-8B: the full encoder behind the small heads, reusable candidates, private APIs and verifier evaluation limits.

Abstract option-score bars illustrating decision calibration, not measured benchmark values AI & Agents

AnyJev: LLM Calibration and Option-Order Bias

A deeper look at Nokia’s AnyJev: why stable answers are not calibrated probabilities, what version 0.2.0 changes and how to evaluate a real router.

Complete article directory

  1. Context Language Models vs Compaction: What to Pilot
  2. LiteLLM Lens: Agent Trace Analysis with SQL and APIs
  3. SmythOS Studio Self-Hosting: Docker Setup, Costs and Limits
  4. DeerFlow 2.0: Docker Setup, Sandboxes and Memory
  5. CLM-8B Self-Hosting: vLLM, Action Cache and Verifiers
  6. AnyJev: LLM Calibration and Option-Order Bias
  7. LiteAgents SDK: Per-Turn Routing, Setup and Migration
  8. mcp-memory-service: Shared Memory for Claude Code and Cursor
  9. Claude Opus 5.5: Best Uses, Prompts and Effort Settings
  10. Laya and Jev for Business: 6 Practical Workflows
  11. Laya vs Jev: What the Benchmarks Mean for AI Startups
  12. Jev AI Review: Decision Models for Agent Workflows
  13. SwarmLLM Review 2026: Browser P2P LLM Inference Across Phones and Laptops
  14. Phonely Alma Review: Is the Voice LLM Ready for Production?
  15. Utopia Review: Temporal Knowledge Graph for Enterprise
  16. NVIDIA PAIR Review: Local AI Routing and the AMD Gap
  17. Tencent Hy4 Preview Review: Is the 1M-Context Coding Model Worth a Pilot?
  18. PageLM Review: Self-Hosting and Commercial Use
  19. M6 Mac mini vs M5 Mac Studio for Local AI: Buyer's Guide
  20. FreeToken AI Review: Run Frontier MoE Models on Consumer GPUs?
  21. Darkbloom AI Review: Private Inference on Idle Macs
  22. Ox Alpha Revealed as GLM-5.3-Flash: What Changed
  23. Pika Audio Models API Pricing: Is SFX Really 20x Cheaper?
  24. Thunder Compute's $13M GPU Virtualization Bet: Enterprise Buyer's Guide
  25. AirLLM on 4 GB VRAM: How Layer-Wise Inference Really Works
  26. Qwen3.8-27B: Self-Hosted Computer-Use Agents Without Exporting Screenshots
  27. Netflix's vLLM and Triton Stack: 7 Production Lessons
  28. Transformers.js Browser AI: When Local Inference Belongs in Your Product
  29. How to Self-Host LiteLLM in Production: 2026 Guide
  30. AI-Ready Company Wiki: Architecture and Build Guide
  31. Does Claude Watermark Text? The 2026 API Answer
  32. OpenKB Review: Knowledge Compiler vs RAG
  33. Unsloth Desktop Review: A Private Local AI Workstation?
  34. NeMo Switchyard 0.2: Agent Model Routing Without Training?
  35. Firecrawl AnyDoc Review: 14 Formats to Markdown
  36. Muse Glimmer 30B: Is Meta's Local Agent Model Production-Ready?
  37. OmniRoute AI Routing: Setup and Production Checklist
  38. Gemini Robotics 2: Whole-Body Control and the Pilot Decision
  39. pdf-inspector Review: Route PDFs Before OCR
  40. Local Multimodal AI Coding Assistant: Voice, OCR and Privacy
  41. DeepSeek V4 Flash 0731 on One AI PC: What Actually Works?
  42. Fine-Tune Gemma 4 Free with Unsloth and Colab
  43. Taalas HC1 Review: Is a Hardwired LLM ASIC Worth It?
  44. llmfit Guide: Which Local LLM Fits Your Hardware?
  45. Miso TTS Self-Hosted vs API: Cost, Latency and VRAM Reality
  46. Compress RAG Vectors to 2 Bits: Is Data-Oblivious Quantization Production-Ready?
  47. LMCache Reported 13.7x Lower Mean TTFT with a Shared KV Cache
  48. Lossless LLM Weight Compression vs 8-bit GGUF: What Is Ready for Production?
  49. Cisco Antares Review: 1B Local AI for Vulnerability Localization
  50. NVIDIA Nemotron 3.5 ASR: Is Free Self-Hosted STT Ready for Voice Agents?
  51. Kimi K3 for EU Companies: API Cost, Data Risk, and a Pilot Plan
  52. Mesh LLM Review: Can One Large LLM Run Across Multiple Computers?
  53. Bonsai 27B Review: Can a 27B LLM Really Run on a Phone?
  54. Soofi S: Is Germany's Sovereign LLM Ready for Business?
  55. Colibri Runs GLM-5.2 on Consumer Hardware. Here Is the Catch.
  56. When Local Models Beat APIs: A Break-Even Calculator for EU Companies
  57. LLM Cost Calculator 2026: Cost per Task, Not Cost per Token
  58. Pxpipe Review: Can Images Cut Claude Code Token Costs 60%?
  59. Cheaper Per Token Can Still Cost More Per Task
  60. Anthropic Says No Open-Weights Ban. The Cost War Is Still Real.
  61. LLM Gateways Compared 2026: LiteLLM vs OpenRouter vs Portkey vs RouteLLM
  62. Self-Hosting LLMs in the EU: When Open Weights Actually Pay Off
  63. Best Open-Weight LLMs 2026: DeepSeek vs Qwen vs Kimi vs GLM vs Llama
  64. How to Cut LLM Token Costs in 2026
  65. When Is an LLM Eval Worth Building? Cost, ROI, and Trusting the Judge
  66. RAG vs Fine-Tuning vs Long-Context 2026
  67. LLM API Costs 2026. Architecture Shift