
Z.ai released GLM-5.3 on August 14, 2026 — a 743-billion parameter Mixture-of-Experts model. While sharing the same base architecture as GLM-5.2, scaled post-training pushes its Terminal-Bench 3.0 performance from 4.6% to 28.3%.
Short, same-day coverage of notable AI model releases and announcements. These briefs give you the core facts — what was released, how it compares, and where to find more — without the padding.

Z.ai released GLM-5.3 on August 14, 2026 — a 743-billion parameter Mixture-of-Experts model. While sharing the same base architecture as GLM-5.2, scaled post-training pushes its Terminal-Bench 3.0 performance from 4.6% to 28.3%.

Google announced the expansion of Google Antigravity to enterprise customers on August 20, 2026. The agentic development platform is now integrated into eligible Gemini Enterprise app subscriptions, offering administrative spend controls and extensions for VS Code.

Google released Gemini 3.7 Flash on August 13, 2026 — a post-training update to its mid-tier workhorse model that doubles DeepSWE scores and nearly triples AutomationBench performance over Gemini 3.6 Flash, while launching at introductory pricing of $0.75/$3.75 per million tokens through the end of 2026.

OpenAI released updates to its Model Spec framework on August 18, 2026. The new guidelines establish specific rules for interactions with teens, clarify how models should address false or unsupported user premises, and mandate transparency regarding AI capabilities and limitations.

Alibaba's Tongyi Lab released Qwen 3.8-27B on August 14, 2026 — a 27.8-billion parameter dense vision-language model under Apache 2.0. Supporting text, image, and video inputs with a native 262K context window, it runs in roughly 16-17 GB of VRAM when quantized to 4-bit, making it viable on a single RTX 4090.

Pokee AI has launched Pokee-Isaac 28B v0, an agentic model claiming a 10-million-token context window that can run entirely within customer boundaries on consumer hardware like a single RTX 4090 GPU. The model aims to eliminate the need for traditional RAG by processing entire codebases locally.