Aakib Ansari.

All Articles

Explore all our published content, including real-world model tests, detailed deep dives, and breaking industry updates.

OpenAI Cuts GPT-5.6 Luna Prices 80% and Terra 20% After Sol Optimises Its Own Inference
News Brief2 min read
OpenAI Cuts GPT-5.6 Luna Prices 80% and Terra 20% After Sol Optimises Its Own Inference

OpenAI cut API prices for GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20% on July 30, citing GPU kernel optimisations and speculative decoding improvements driven by Sol. Luna drops to $0.20/$1.20 per million tokens. Terra drops to $2.00/$12.00. Sol's standard price is unchanged, but a new Fast mode — 2.5× throughput at double the rate — replaces the old Priority Processing tier.

MiniMax H3 Generates 2K Video With Native Stereo Audio in One Shot
News Brief2 min read
MiniMax H3 Generates 2K Video With Native Stereo Audio in One Shot

MiniMax launched H3 (also known as Hailuo H3) on July 31, a general-purpose multimodal model that generates 2K video at 24fps with native, synchronised stereo audio in a single pass. Previous MiniMax models were specialised video generators; H3 unifies text, image, video, and audio into one context. Open weights are planned. The model is live via the Hailuo AI app and the MiniMax Open Platform API.

Mira Murati's Thinking Machines Ships Inkling-Small: 276B Open-Weight Model Within One Point of Its Flagship
News Brief2 min read
Mira Murati's Thinking Machines Ships Inkling-Small: 276B Open-Weight Model Within One Point of Its Flagship

Thinking Machines Lab, the startup founded by former OpenAI CTO Mira Murati, released Inkling-Small on July 30 — two weeks after the 975B flagship Inkling. The small model has 276B total and 12B active parameters, runs under an Apache 2.0 licence, and scores 40 on the Artificial Analysis Intelligence Index, within one point of the flagship's 41. It is designed for enterprise fine-tuning via the lab's Tinker platform.

DeepSeek-V4-Flash-0731: The Agentic Coding Upgrade That Closes the Gap With Pro
News Brief2 min read
DeepSeek-V4-Flash-0731: The Agentic Coding Upgrade That Closes the Gap With Pro

DeepSeek released V4-Flash-0731 on July 31, an updated version of their 284B MoE Flash model re-post-trained specifically for agentic and coding performance. TerminalBench 82.7 and NL2Repo 54.2 are the headline numbers. The Artificial Analysis Intelligence Index puts it 10 points above the previous Flash and competitively close to Pro-tier models. MIT-licensed weights and GGUF versions are already available.

Alibaba's Qwen 3.7 Flash Is the $0.03 Per Million Token Multimodal Workhorse
News Brief2 min read
Alibaba's Qwen 3.7 Flash Is the $0.03 Per Million Token Multimodal Workhorse

Alibaba released Qwen 3.7 Flash on July 27, a cost-optimised multimodal model priced at $0.03 input / $0.13 output per million tokens with a 1M context window. It accepts text, image, and video input and is designed for high-volume agent loops, browser automation, and support tooling. It is the flash-tier counterpart to the Qwen 3.7 Max, Alibaba's flagship agentic model released in May.