Sapiens AI Launches Agnes 2.5 Pro Beta: $0.10/M 1M-Context Reasoning Model


Singapore-based artificial intelligence developer Sapiens AI officially launched Agnes 2.5 Pro Beta on August 26, 2026, introducing an aggressive low-cost reasoning model designed for high-throughput enterprise pipelines.
Agnes 2.5 Pro Beta features a 1-million-token context window (roughly 1,500 pages of text) paired with a 65,536-token maximum generation limit per request. In independent benchmarking published by Artificial Analysis, the model achieved an Intelligence Index score of 49—placing it in competitive proximity to OpenAI's GPT-5.6 Luna (50)—while sustaining a generation throughput of 159.5 tokens per second.
The model's primary market disruption lies in its commercial pricing structure. Sapiens AI set API access rates at $0.10 per million input tokens and $0.40 per million output tokens, accompanied by a 90% discount on cached prompt inputs ($0.01 per million cached tokens). The model supports multimodal inputs via image URLs and exposes native endpoints conforming to standard OpenAI Chat Completions and Anthropic Messages protocols.
By pairing frontier-tier context lengths with budget-tier inference economics, Agnes 2.5 Pro Beta directly targets agentic loops and massive document batching workflows previously dominated by Google's Gemini 3.5 Flash-Lite and Z.AI's [GLM-5.2 Turbo](/articles/glm-5-2-turbo-release).


