Models and infrastructure
Model selection, inference economics, local deployment, compression and serving architecture.
- Blog posts
- 46
- Topic
- AI and agents
Start with the cornerstone
Latest in this collection
Ox Alpha Free AI Model: Setup, Privacy and Buyer Guide
Use the 1M-context stealth coding model through OpenCode or OpenRouter. Compare IDs, privacy terms, hidden costs and a safe team evaluation plan.
Pika Audio Models API Pricing: Is SFX Really 20x Cheaper?
Pika lists SFX at $0.0002 per second, plus three more audio models. See the real break-even, API limits, quality gates and production fit before you switch.
Thunder Compute's $13M GPU Virtualization Bet: Enterprise Buyer's Guide
Thunder Compute says GPU pooling can recover stranded capacity. Compare network virtualization with MIG, vGPU and passthrough, then scope a measurable enterprise pilot.
AirLLM on 4 GB VRAM: How Layer-Wise Inference Really Works
AirLLM can execute models far larger than GPU memory by streaming weights layer by layer. Separate the measured VRAM claim from disk, latency, quantization and production reality.
Qwen3.8-27B: Self-Hosted Computer-Use Agents Without Exporting Screenshots
An open-weight model now leads desktop-operation benchmarks. What that changes for back-office automation you cannot send to a hosted API.
Netflix's vLLM and Triton Stack: 7 Production Lessons
Netflix chose vLLM for operational fit, then found the real bottlenecks in constrained decoding, deployment, model loading and metrics. Here is what smaller teams should copy, and what they should buy instead.
Complete article directory
- Ox Alpha Free AI Model: Setup, Privacy and Buyer Guide
- Pika Audio Models API Pricing: Is SFX Really 20x Cheaper?
- Thunder Compute's $13M GPU Virtualization Bet: Enterprise Buyer's Guide
- AirLLM on 4 GB VRAM: How Layer-Wise Inference Really Works
- Qwen3.8-27B: Self-Hosted Computer-Use Agents Without Exporting Screenshots
- Netflix's vLLM and Triton Stack: 7 Production Lessons
- Transformers.js Browser AI: When Local Inference Belongs in Your Product
- How to Self-Host LiteLLM in Production: 2026 Guide
- AI-Ready Company Wiki: Architecture and Build Guide
- Does Claude Watermark Text? The 2026 API Answer
- OpenKB Review: Knowledge Compiler vs RAG
- Unsloth Desktop Review: A Private Local AI Workstation?
- NeMo Switchyard 0.2: Agent Model Routing Without Training?
- Firecrawl AnyDoc Review: 14 Formats to Markdown
- Muse Glimmer 30B: Is Meta's Local Agent Model Production-Ready?
- OmniRoute AI Routing: Setup and Production Checklist
- Gemini Robotics 2: Whole-Body Control and the Pilot Decision
- pdf-inspector Review: Route PDFs Before OCR
- Local Multimodal AI Coding Assistant: Voice, OCR and Privacy
- DeepSeek V4 Flash 0731 on One AI PC: What Actually Works?
- Fine-Tune Gemma 4 Free with Unsloth and Colab
- Taalas HC1 Review: Is a Hardwired LLM ASIC Worth It?
- llmfit Guide: Which Local LLM Fits Your Hardware?
- Miso TTS Self-Hosted vs API: Cost, Latency and VRAM Reality
- Cut RAG Vector Memory 16x: Is Data-Oblivious Quantization Ready for Production?
- Shared KV Cache Cut LLM Inference Latency 14x, With No New GPUs
- Lossless LLM Weight Compression vs 8-bit GGUF: What Is Ready for Production?
- Cisco Antares Review: 1B Local AI for Vulnerability Localization
- NVIDIA Nemotron 3.5 ASR: Is Free Self-Hosted STT Ready for Voice Agents?
- Kimi K3 for EU Companies: API Cost, Data Risk, and a Pilot Plan
- Mesh LLM Review: Can One Large LLM Run Across Multiple Computers?
- Bonsai 27B Review: Can a 27B LLM Really Run on a Phone?
- Soofi S: Is Germany's Sovereign LLM Ready for Business?
- Colibri Runs GLM-5.2 on Consumer Hardware. Here Is the Catch.
- When Local Models Beat APIs: A Break-Even Calculator for EU Companies
- LLM Cost Calculator 2026: Cost per Task, Not Cost per Token
- Pxpipe Review: Can Images Cut Claude Code Token Costs 60%?
- Cheaper Per Token. More Expensive Per Answer.
- Dario Declared War on Open Source. The Real War Is Over Your AI Bill.
- LLM Gateways Compared 2026: LiteLLM vs OpenRouter vs Portkey vs RouteLLM
- Self-Hosting LLMs in the EU: When Open Weights Actually Pay Off
- Best Open-Weight LLMs 2026: DeepSeek vs Qwen vs Kimi vs GLM vs Llama
- How to Cut LLM Token Costs in 2026
- When Is an LLM Eval Worth Building? Cost, ROI, and Trusting the Judge
- RAG vs Fine-Tuning vs Long-Context 2026
- LLM API Costs 2026. Architecture Shift