THE NIELLO JOURNAL
Ideas worth
connecting.
Guides, perspectives and technology notes on putting AI to work in the enterprise. 83 articles, newest first.
DeepSeek-V4.1-Flash: MIT Weights, 552B Parameters, and the VRAM Maths
An MIT-licensed model just topped Terminal-Bench 2.1. The architecture, the serving arithmetic, and when to self-host.
Read the article ↗82 articles
GPT-6 Astra: API Pricing, the 272K Cliff, and What the Benchmarks Actually Say
Astra costs $10/$50 per million tokens — but the 272K billing cliff and the cost-per-task figures matter more.
AI Image and Video Models in 2026: A Buyer Comparison With Real Per-Shot Costs
Veo 3.1, Kling 3.0, Seedance and Flux 2 compared on price, licensing and cost per usable shot.
Qwen3.8-Max: Alibaba Ships 2.4 Trillion Parameters, With Open Weights Promised
Qwen3.8-Max launched at $2/$6 per million tokens. What is verified, what is vendor claim, and what is missing.
How Cheap Is DeepSeek, Really? A Cost Comparison Against Every Leading Model
DeepSeek V4-Pro costs $0.435/$0.87 per million tokens. The real multiple against GPT-5.6, Fable 5 and Gemini.
Where Token Prices Are Going: What to Budget for LLM Costs Through 2027
LLM token prices keep falling while AI bills keep rising. The four forces driving each, and how to budget for both.
LLM Model Distillation Explained: Why Distillation-as-a-Service Is the Missing Layer
Distillation turns a model nobody can run into one anyone can. The mechanics, economics, and why we explore it.
Kimi K3 Open Weights Are Out: Download, Self-Hosting Requirements, and What Changed
Moonshot shipped 2.8T parameters on schedule. Independent numbers, serving math, and why policy cannot undo it.
How to Run a 27B LLM on Your Phone: Bonsai 27B, 1-Bit and Ternary Quantization
Bonsai 27B fits a 27B model in 3.9 GB. What survives 1-bit and ternary compression, and what quietly breaks.
Kimi K3 and Inkling: The Week Open Weights Reached the Frontier — and What It Takes to Run Them
Moonshot's 2.8T Kimi K3 and Murati's Inkling landed a day apart: the power, the GPU reality, July 2026's model wave.
The Efficiency Race: Why AI Now Competes on Cost per Intelligence
AI is shifting from raw IQ to cost per unit of intelligence. Why efficiency, not size, is the race that matters now.
Reasoning in a Private Language: From Chain-of-Thought to Latent Thinking
How models are learning to think in a compressed private notation instead of English — and why that quietly lowers cost.
The Hidden Reasoning-Token Tax in LLM Pricing (and How GPT-5.6 Cut It)
Reasoning models bill you for hidden thinking tokens you never see. Here's the real cost, and how GPT-5.6 cut it.