Aakib Ansari.
Back to articles
News Brief

Sapiens AI Launches Agnes 2.5 Pro Beta: $0.10/M 1M-Context Reasoning Model

Md Aakib Ansari
Md Aakib AnsariWeb Developer & AI Tools Reviewer
•3 min read•Model: Sapiens AI Agnes 2.5 Pro Beta•Company: Sapiens AI
Sapiens AI Launches Agnes 2.5 Pro Beta: $0.10/M 1M-Context Reasoning Model

Singapore-based artificial intelligence developer Sapiens AI officially launched Agnes 2.5 Pro Beta on August 26, 2026, introducing an aggressive low-cost reasoning model designed for high-throughput enterprise pipelines.

Agnes 2.5 Pro Beta features a 1-million-token context window (roughly 1,500 pages of text) paired with a 65,536-token maximum generation limit per request. In independent benchmarking published by Artificial Analysis, the model achieved an Intelligence Index score of 49—placing it in competitive proximity to OpenAI's GPT-5.6 Luna (50)—while sustaining a generation throughput of 159.5 tokens per second.

The model's primary market disruption lies in its commercial pricing structure. Sapiens AI set API access rates at $0.10 per million input tokens and $0.40 per million output tokens, accompanied by a 90% discount on cached prompt inputs ($0.01 per million cached tokens). The model supports multimodal inputs via image URLs and exposes native endpoints conforming to standard OpenAI Chat Completions and Anthropic Messages protocols.

By pairing frontier-tier context lengths with budget-tier inference economics, Agnes 2.5 Pro Beta directly targets agentic loops and massive document batching workflows previously dominated by Google's Gemini 3.5 Flash-Lite and Z.AI's [GLM-5.2 Turbo](/articles/glm-5-2-turbo-release).

Frequently Asked Questions

What is Sapiens AI Agnes 2.5 Pro Beta?
Agnes 2.5 Pro Beta is a proprietary 1M-context reasoning model developed by Singapore's Sapiens AI, optimized for low-latency token throughput and high-volume document ingestion.
How much does the Agnes 2.5 Pro Beta API cost?
The model is priced at $0.10 per million input tokens ($0.01/M for cached prompts) and $0.40 per million output tokens, supporting up to 65,536 output tokens per request.

Related Articles