FROM THE TRENCHES Issue №164

Opinions, Not Press Releases

Browse a curated front page, then move through focused topic and cluster hubs without losing older work to chronology.

Explore by topic
06
Focused clusters
12
Blog posts
164

Latest articles

Nine articles per page, ordered by publication date.

Floci, LocalStack and AWS SAM compared for local AWS integration testing Delivery & QA

Floci vs LocalStack: AWS Emulator Review for CI in 2026

A fact-checked buyer guide to Floci, LocalStack, AWS SAM and real AWS: fidelity, service coverage, licensing, CI fit and a low-risk migration test.

A voice waveform splitting between a managed Miso TTS API and a private self-hosted GPU AI & Agents

Miso TTS Self-Hosted vs API: Cost, Latency and VRAM Reality

Miso TTS is expressive and open-weight, but the 110 ms claim belongs to its hosted H100 API, not local inference. Compare hardware, pricing, licensing and production risk before choosing.

An MCP client presents an audience-bound token to a resource server that validates it against an identity provider before any tool runs AI & Agents

Enterprise MCP Authorization Architecture

OAuth 2.1, audience-bound tokens and no token passthrough are the spec's hard rules. The rest, multi-tenant isolation, per-tool scopes, delegated access and audit, is a reference design. Here is the whole trust boundary in two diagrams.

EU AI Act Article 50 transparency obligations mapped to a build checklist for SaaS chatbots, AI agents and generated content Business & Regulation

EU AI Act Article 50 Checklist for SaaS and AI Agents

Article 50 applies 2 August 2026. The engineering build checklist: chatbot and agent disclosure, C2PA machine-readable marking, deepfake labels, public-interest text, logging, and the evidence to keep for an audit.

Data-oblivious vector quantization compressing a dense float32 embedding field into a compact 2-bit bucket grid at 16x smaller memory AI & Agents

Cut RAG Vector Memory 16x: Is Data-Oblivious Quantization Ready for Production?

A new Rust index shrinks 10M embeddings from 31 GB to 4 GB with Google's training-free TurboQuant. What the method proves, where it beats FAISS, the recall trade-off, and how to pilot it.

Eight isolated per-rank KV caches on one server collapsing into a single shared host-side cache layer that every serving process can read AI & Agents

Shared KV Cache Cut LLM Inference Latency 14x, With No New GPUs

An open-source KV cache layer cut mean time-to-first-token from 3.98s to 0.29s on a 235B model. Same GPUs, same memory. It stopped eight processes from hoarding private caches. What changed and whether it helps you.

Lossless BF16 weight compression compared with 8-bit GGUF quantization across exactness, runtime evidence and production readiness AI & Agents

Lossless LLM Weight Compression vs 8-bit GGUF: What Is Ready for Production?

A fact-checked decision guide to the GLM-5.2 result, exact BF16 decoding, Q8 quantization, prior systems, runtime evidence, fleet economics and a 14-day pilot.

Paid pilot, proof of concept and design partner compared by the commercial evidence each produces Product & MVP

Paid Pilot vs PoC vs Design Partner: Which One Proves Demand?

Decision guide for choosing the right evidence test: technical feasibility, customer learning or willingness to pay, with a seven-condition paid-pilot scorecard and contract brief.

Nested AI agent evaluation sandbox with isolated compute, controlled egress, ephemeral credentials and out-of-band monitoring AI & Agents

OpenAI Eval Sandbox Escape: 12 Controls Before You Test a Cyber Agent

A fact-checked incident analysis and procurement checklist for nested isolation, egress, credentials, control-plane separation, benchmark integrity, monitoring and forensic fallback.

Inbox, without the noise

Follow the work that matters to you

Get a short email when we publish something new. Follow the whole blog or only the problems you care about.

What would you like to receive?
Choose your topics

Free, double opt-in, no tracking pixels.