
IBM has released Granite 4.2, a family of dense open-weight reasoning models in 3B, 8B, and 30B sizes under the Apache 2.0 license, featuring switchable thinking mode, 512K context, and agentic RL.
Short, same-day coverage of notable AI model releases and announcements. These briefs give you the core facts — what was released, how it compares, and where to find more — without the padding.

IBM has released Granite 4.2, a family of dense open-weight reasoning models in 3B, 8B, and 30B sizes under the Apache 2.0 license, featuring switchable thinking mode, 512K context, and agentic RL.

Alibaba has released Qwen3.8-Flash-Next, an open-weight 125B multimodal MoE model activating only 6B parameters per token under the Qwen Community 1.0 license, previewing the next-generation Qwen4 architecture.

Z.AI has launched GLM-5.2 Turbo, an inference-optimized 744B open-weight MoE model with a 1M token context window under the MIT license, delivering lower latency for high-frequency coding agent workflows.

An unannounced AI model named Ox Alpha has appeared on OpenRouter and OpenCode with a 1M token context window, multimodal capabilities, and free tier access.

Snowflake announced a major update to its Cortex AI Gateway on August 18, 2026, introducing dynamic model routing. The new feature automatically directs enterprise prompts to the most cost-effective LLM based on task complexity, reducing API expenses by up to 50%.

Alibaba released the open-weight checkpoint for Qwen 3.8-Max on August 12-13, 2026 under a custom license. Designated Qwen3.8-2.4T-A95B, the weights cover the full 2.4-trillion-parameter MoE architecture with 95B active parameters — but the public checkpoint is text-only, without the multimodal and 1M-context capabilities of the API version.