Cluster

Models and infrastructure

Model selection, inference economics, local deployment, compression and serving architecture.

Blog posts
46
Topic
AI and agents

Start with the cornerstone

Latest in this collection

Ox Alpha free stealth AI model routed through OpenCode Zen and OpenRouter with privacy and cost checks AI & Agents

Ox Alpha Free AI Model: Setup, Privacy and Buyer Guide

Use the 1M-context stealth coding model through OpenCode or OpenRouter. Compare IDs, privacy terms, hidden costs and a safe team evaluation plan.

Four Pika Audio API models mapped to sound effects, speech, music and synchronized video soundtracks with cost controls AI & Agents

Pika Audio Models API Pricing: Is SFX Really 20x Cheaper?

Pika lists SFX at $0.0002 per second, plus three more audio models. See the real break-even, API limits, quality gates and production fit before you switch.

GPU workloads sharing a pooled accelerator fleet through a virtualization layer AI & Agents

Thunder Compute's $13M GPU Virtualization Bet: Enterprise Buyer's Guide

Thunder Compute says GPU pooling can recover stranded capacity. Compare network virtualization with MIG, vGPU and passthrough, then scope a measurable enterprise pilot.

AirLLM streams one model layer from storage through RAM to limited GPU VRAM AI & Agents

AirLLM on 4 GB VRAM: How Layer-Wise Inference Really Works

AirLLM can execute models far larger than GPU memory by streaming weights layer by layer. Separate the measured VRAM claim from disk, latency, quantization and production reality.

Self-hosted screen agent operating a desktop interface inside a single-GPU deployment AI & Agents

Qwen3.8-27B: Self-Hosted Computer-Use Agents Without Exporting Screenshots

An open-weight model now leads desktop-operation benchmarks. What that changes for back-office automation you cannot send to a hosted API.

Layered vLLM and NVIDIA Triton inference stack AI & Agents

Netflix's vLLM and Triton Stack: 7 Production Lessons

Netflix chose vLLM for operational fit, then found the real bottlenecks in constrained decoding, deployment, model loading and metrics. Here is what smaller teams should copy, and what they should buy instead.

Complete article directory

  1. Ox Alpha Free AI Model: Setup, Privacy and Buyer Guide
  2. Pika Audio Models API Pricing: Is SFX Really 20x Cheaper?
  3. Thunder Compute's $13M GPU Virtualization Bet: Enterprise Buyer's Guide
  4. AirLLM on 4 GB VRAM: How Layer-Wise Inference Really Works
  5. Qwen3.8-27B: Self-Hosted Computer-Use Agents Without Exporting Screenshots
  6. Netflix's vLLM and Triton Stack: 7 Production Lessons
  7. Transformers.js Browser AI: When Local Inference Belongs in Your Product
  8. How to Self-Host LiteLLM in Production: 2026 Guide
  9. AI-Ready Company Wiki: Architecture and Build Guide
  10. Does Claude Watermark Text? The 2026 API Answer
  11. OpenKB Review: Knowledge Compiler vs RAG
  12. Unsloth Desktop Review: A Private Local AI Workstation?
  13. NeMo Switchyard 0.2: Agent Model Routing Without Training?
  14. Firecrawl AnyDoc Review: 14 Formats to Markdown
  15. Muse Glimmer 30B: Is Meta's Local Agent Model Production-Ready?
  16. OmniRoute AI Routing: Setup and Production Checklist
  17. Gemini Robotics 2: Whole-Body Control and the Pilot Decision
  18. pdf-inspector Review: Route PDFs Before OCR
  19. Local Multimodal AI Coding Assistant: Voice, OCR and Privacy
  20. DeepSeek V4 Flash 0731 on One AI PC: What Actually Works?
  21. Fine-Tune Gemma 4 Free with Unsloth and Colab
  22. Taalas HC1 Review: Is a Hardwired LLM ASIC Worth It?
  23. llmfit Guide: Which Local LLM Fits Your Hardware?
  24. Miso TTS Self-Hosted vs API: Cost, Latency and VRAM Reality
  25. Cut RAG Vector Memory 16x: Is Data-Oblivious Quantization Ready for Production?
  26. Shared KV Cache Cut LLM Inference Latency 14x, With No New GPUs
  27. Lossless LLM Weight Compression vs 8-bit GGUF: What Is Ready for Production?
  28. Cisco Antares Review: 1B Local AI for Vulnerability Localization
  29. NVIDIA Nemotron 3.5 ASR: Is Free Self-Hosted STT Ready for Voice Agents?
  30. Kimi K3 for EU Companies: API Cost, Data Risk, and a Pilot Plan
  31. Mesh LLM Review: Can One Large LLM Run Across Multiple Computers?
  32. Bonsai 27B Review: Can a 27B LLM Really Run on a Phone?
  33. Soofi S: Is Germany's Sovereign LLM Ready for Business?
  34. Colibri Runs GLM-5.2 on Consumer Hardware. Here Is the Catch.
  35. When Local Models Beat APIs: A Break-Even Calculator for EU Companies
  36. LLM Cost Calculator 2026: Cost per Task, Not Cost per Token
  37. Pxpipe Review: Can Images Cut Claude Code Token Costs 60%?
  38. Cheaper Per Token. More Expensive Per Answer.
  39. Dario Declared War on Open Source. The Real War Is Over Your AI Bill.
  40. LLM Gateways Compared 2026: LiteLLM vs OpenRouter vs Portkey vs RouteLLM
  41. Self-Hosting LLMs in the EU: When Open Weights Actually Pay Off
  42. Best Open-Weight LLMs 2026: DeepSeek vs Qwen vs Kimi vs GLM vs Llama
  43. How to Cut LLM Token Costs in 2026
  44. When Is an LLM Eval Worth Building? Cost, ROI, and Trusting the Judge
  45. RAG vs Fine-Tuning vs Long-Context 2026
  46. LLM API Costs 2026. Architecture Shift