FROM THE TRENCHES Issue №164

Opinions, Not Press Releases

Browse a curated front page, then move through focused topic and cluster hubs without losing older work to chronology.

Explore by topic
06
Focused clusters
12
Blog posts
164

Latest articles

Nine articles per page, ordered by publication date.

EU AI vendor security questionnaire with 45 evidence and contract checks Business & Regulation

EU AI Vendor Security Questionnaire: 45 Questions Before You Sign

An evidence-led AI procurement checklist covering training, retention, subprocessors, regions, logs, permissions, model changes, incidents, evals, deletion, exit and EU AI Act roles. Includes a multilingual spreadsheet.

DACH AI adoption benchmark comparing Austria, Germany and Switzerland in 2026 Business & Regulation

DACH AI Adoption Benchmark 2026: What SMEs Put Into Production

Source-backed adoption and use-case data for Austria, Germany and Switzerland, plus the production, budget, ownership and shelfware metrics public surveys still do not measure.

Graphify codebase knowledge graph connecting code, data, infrastructure and documentation AI & Agents

Graphify Review 2026: Is a Codebase Knowledge Graph Worth It?

Buyer review of Graphify vs search and RAG: architecture, privacy boundaries, benchmark limits, real adoption cost and a two-week pilot scorecard.

Bonsai 27B 1-bit local AI model compressed from 54 GB to a phone-class footprint AI & Agents

Bonsai 27B Review: Can a 27B LLM Really Run on a Phone?

Verified buyer's review of 1-bit vs ternary, real memory, uneven benchmark losses, phone and WebGPU speed, use cases and pilot gates.

Soofi S European sovereign LLM procurement review AI & Agents

Soofi S: Is Germany's Sovereign LLM Ready for Business?

Germany's 31.6B sparse model is strong on German and code, but the current checkpoint is a selected-partner closed beta with no final license. Benchmarks, openness, infrastructure, alternatives and the pilot checklist EU buyers need.

AI pilot kill-or-scale scorecard with 12 business, quality, reliability and adoption metrics AI & Agents

AI Pilot Kill-or-Scale Scorecard: 12 Metrics to Check After 30 Days

A decision-grade scorecard with formulas, hard gates and one worked example across baseline cost, successful actions, straight-through completion, correction time, failures, latency, adoption, auditability, data readiness and payback.

T3MP3ST autonomous AI red-team harness reviewed against OWASP APTS AI & Agents

T3MP3ST Review 2026: Can It Replace a Penetration Test?

A buyer-focused review separating benchmarked single-agent capabilities from the unproven swarm, mapped against OWASP APTS with a safe pilot decision for CTOs.

Vibe coder and junior developer taking different paths through an AI-assisted engineering system Leadership & Teams

Are Vibe Coders the New Junior Developers?

The traditional ticket-taking junior is shrinking, but the junior developer is not extinct. Current labour data, the difference between prompting and engineering, and a hiring model for AI-native entry-level talent.

External QA benchmark for software defects found in the first 30 days Delivery & QA

External QA Benchmark: What We Find in the First 30 Days

A research-backed defect benchmark for permissions, regressions, browser/device combinations, data integrity, AI failures and production escapes, and why no honest universal bug-count median exists.

Inbox, without the noise

Follow the work that matters to you

Get a short email when we publish something new. Follow the whole blog or only the problems you care about.

What would you like to receive?
Choose your topics

Free, double opt-in, no tracking pixels.