Cluster

Models and infrastructure

Model selection, inference economics, local deployment, compression and serving architecture.

Blog posts
55
Topic
AI and agents

Start with the cornerstone

Inbox, without the noise

Get the next AI and agents field note

One concise email when we publish. No tracking pixels, and no inbox filler.

Latest in this collection

A phone and laptop split one large language model across browser tabs while WebRTC carries activations between them AI & Agents

SwarmLLM Review 2026: Browser P2P LLM Inference Across Phones and Laptops

A buyer-focused review of SwarmLLM: browser-native 27B inference across phones and laptops, WebGPU/WebRTC architecture, benchmark limits, peer trust, and when to pilot it.

A phone waveform passes through the Alma voice LLM into a production evaluation scorecard AI & Agents

Phonely Alma Review: Is the Voice LLM Ready for Production?

Alma claims 182 ms TTFT, a 206 ms P99 and $0.55 per blended million tokens. See what the benchmark proves, what it omits and how to run a buyer-side voice-agent pilot.

Layered Utopia knowledge records represent preserved history AI & Agents

Utopia Review: Temporal Knowledge Graph for Enterprise

Evaluate Utopia for enterprise knowledge: two-clock history, source provenance, self-hosting limits and a practical pilot plan before you invest.

NVIDIA PAIR routes local AI agent requests across RTX, Mac and an AMD Strix Halo node AI & Agents

NVIDIA PAIR Review: Local AI Routing and the AMD Gap

PAIR routes independent agent requests across local computers without pooling GPUs. See how it works, where Strix Halo falls short and what an AMD port must prove.

Tencent Hy4 preview routes a million-token codebase through coding-agent tools and verification gates AI & Agents

Tencent Hy4 Preview Review: Is the 1M-Context Coding Model Worth a Pilot?

A buyer-focused Hy4 preview review covering coding benchmarks, WorkBuddy, API cost, self-hosting hardware, known limits and a two-week evaluation plan.

PageLM document learning platform: self-hosting, licensing and data-flow review AI & Agents

PageLM Review: Self-Hosting and Commercial Use

A source-checked PageLM review for teams: what it generates, where data goes, what commercial use requires and how to test a training pilot.

Complete article directory

  1. SwarmLLM Review 2026: Browser P2P LLM Inference Across Phones and Laptops
  2. Phonely Alma Review: Is the Voice LLM Ready for Production?
  3. Utopia Review: Temporal Knowledge Graph for Enterprise
  4. NVIDIA PAIR Review: Local AI Routing and the AMD Gap
  5. Tencent Hy4 Preview Review: Is the 1M-Context Coding Model Worth a Pilot?
  6. PageLM Review: Self-Hosting and Commercial Use
  7. M6 Mac mini vs M5 Mac Studio for Local AI: Buyer's Guide
  8. FreeToken AI Review: Run Frontier MoE Models on Consumer GPUs?
  9. Darkbloom AI Review: Private Inference on Idle Macs
  10. Ox Alpha Revealed as GLM-5.3-Flash: What Changed
  11. Pika Audio Models API Pricing: Is SFX Really 20x Cheaper?
  12. Thunder Compute's $13M GPU Virtualization Bet: Enterprise Buyer's Guide
  13. AirLLM on 4 GB VRAM: How Layer-Wise Inference Really Works
  14. Qwen3.8-27B: Self-Hosted Computer-Use Agents Without Exporting Screenshots
  15. Netflix's vLLM and Triton Stack: 7 Production Lessons
  16. Transformers.js Browser AI: When Local Inference Belongs in Your Product
  17. How to Self-Host LiteLLM in Production: 2026 Guide
  18. AI-Ready Company Wiki: Architecture and Build Guide
  19. Does Claude Watermark Text? The 2026 API Answer
  20. OpenKB Review: Knowledge Compiler vs RAG
  21. Unsloth Desktop Review: A Private Local AI Workstation?
  22. NeMo Switchyard 0.2: Agent Model Routing Without Training?
  23. Firecrawl AnyDoc Review: 14 Formats to Markdown
  24. Muse Glimmer 30B: Is Meta's Local Agent Model Production-Ready?
  25. OmniRoute AI Routing: Setup and Production Checklist
  26. Gemini Robotics 2: Whole-Body Control and the Pilot Decision
  27. pdf-inspector Review: Route PDFs Before OCR
  28. Local Multimodal AI Coding Assistant: Voice, OCR and Privacy
  29. DeepSeek V4 Flash 0731 on One AI PC: What Actually Works?
  30. Fine-Tune Gemma 4 Free with Unsloth and Colab
  31. Taalas HC1 Review: Is a Hardwired LLM ASIC Worth It?
  32. llmfit Guide: Which Local LLM Fits Your Hardware?
  33. Miso TTS Self-Hosted vs API: Cost, Latency and VRAM Reality
  34. Compress RAG Vectors to 2 Bits: Is Data-Oblivious Quantization Production-Ready?
  35. LMCache Reported 13.7x Lower Mean TTFT with a Shared KV Cache
  36. Lossless LLM Weight Compression vs 8-bit GGUF: What Is Ready for Production?
  37. Cisco Antares Review: 1B Local AI for Vulnerability Localization
  38. NVIDIA Nemotron 3.5 ASR: Is Free Self-Hosted STT Ready for Voice Agents?
  39. Kimi K3 for EU Companies: API Cost, Data Risk, and a Pilot Plan
  40. Mesh LLM Review: Can One Large LLM Run Across Multiple Computers?
  41. Bonsai 27B Review: Can a 27B LLM Really Run on a Phone?
  42. Soofi S: Is Germany's Sovereign LLM Ready for Business?
  43. Colibri Runs GLM-5.2 on Consumer Hardware. Here Is the Catch.
  44. When Local Models Beat APIs: A Break-Even Calculator for EU Companies
  45. LLM Cost Calculator 2026: Cost per Task, Not Cost per Token
  46. Pxpipe Review: Can Images Cut Claude Code Token Costs 60%?
  47. Cheaper Per Token Can Still Cost More Per Task
  48. Anthropic Says No Open-Weights Ban. The Cost War Is Still Real.
  49. LLM Gateways Compared 2026: LiteLLM vs OpenRouter vs Portkey vs RouteLLM
  50. Self-Hosting LLMs in the EU: When Open Weights Actually Pay Off
  51. Best Open-Weight LLMs 2026: DeepSeek vs Qwen vs Kimi vs GLM vs Llama
  52. How to Cut LLM Token Costs in 2026
  53. When Is an LLM Eval Worth Building? Cost, ROI, and Trusting the Judge
  54. RAG vs Fine-Tuning vs Long-Context 2026
  55. LLM API Costs 2026. Architecture Shift