Aakib Ansari.
Back to articles
News Brief

Pokee AI Releases Pokee-Isaac 28B v0 With 10-Million Token Context on a Single RTX 4090

Md Aakib Ansari
Md Aakib AnsariWeb Developer & AI Tools Reviewer
2 min readModel: Pokee-Isaac 28B v0
Pokee AI Releases Pokee-Isaac 28B v0 With 10-Million Token Context on a Single RTX 4090

Pokee AI released Pokee-Isaac 28B v0 on August 3, 2026—an agentic large language model designed to handle massive datasets locally. The model is notable for claiming a 10-million-token context window while remaining deployable within customer boundaries on consumer-grade hardware, specifically a single NVIDIA RTX 4090 GPU.

By supporting a 10M-token capacity, Pokee-Isaac 28B is built to digest entire software repositories, multi-year email chains, or massive corporate documentation sets in a single pass, bypassing the need for complex Retrieval-Augmented Generation (RAG) pipelines. According to Pokee AI's technical report, the model achieved a 93.3% performance score on the RULER benchmark at its maximum 10-million-token limit, showing high retrieval accuracy across long sequences, alongside strong results on the Berkeley Function Calling Benchmark (BFCL v4).

While the technical results are promising, some industry observers have pointed out that Pokee AI has not fully disclosed the specifics of its memory compression and attention mechanisms. However, the ability to run such long-context models privately on-device represents a significant development for industries like finance and healthcare that cannot utilize cloud-based APIs due to strict data compliance. The launch adds momentum to the local, privacy-first agent trend seen in other recent releases like Liquid AI's LFM2.5 and inclusionAI's Ling 3.0 Flash.

Frequently Asked Questions

What is Pokee-Isaac 28B v0?
Pokee-Isaac 28B v0 is an agentic AI model developed by Pokee AI that supports a 10-million-token context window and is designed to run locally on consumer hardware.
What hardware is required to run Pokee-Isaac 28B locally?
The model is optimized to run on consumer-grade hardware, including a single NVIDIA RTX 4090 GPU, using advanced memory optimization techniques.
How does it perform on long-context benchmarks?
Pokee AI reports a 93.3% score on the RULER benchmark at 10 million tokens, demonstrating high accuracy in needle-in-a-haystack retrieval tasks.

Related Articles

NVIDIA Nemotron 3.5 Lightning Launches in OCI for Always-On Enterprise Agents
News Brief2 min read
NVIDIA Nemotron 3.5 Lightning Launches in OCI for Always-On Enterprise Agents

NVIDIA has released Nemotron 3.5 Lightning, a model optimized for high-volume, always-on AI agent workflows. Now available within Oracle Cloud Infrastructure (OCI) Enterprise AI, the model is engineered for rapid inference and ultra-low latency, lowering the cost of persistent enterprise automation.

Meta Releases Muse Glimmer: Apache 2.0 Licensed 30B Local Agent Model
News Brief2 min read
Meta Releases Muse Glimmer: Apache 2.0 Licensed 30B Local Agent Model

Meta Superintelligence Labs has released Muse Glimmer, a 30-billion parameter open-weight model distilled from its proprietary Muse Spark flagship. Published under an Apache 2.0 license, Glimmer is purpose-built for offline, on-device agentic workloads like coding, debugging, and file management on consumer hardware.

Ant Group's inclusionAI Team Releases Ling 3.0 Flash FP8 Under MIT License
News Brief2 min read
Ant Group's inclusionAI Team Releases Ling 3.0 Flash FP8 Under MIT License

Ant Group's inclusionAI team has released Ling 3.0 Flash FP8, a highly efficient 124-billion parameter Mixture-of-Experts (MoE) model. Featuring an MIT license and a custom hybrid attention architecture, the model reduces active parameters to 5.1 billion per token, matching the performance of much larger models while dramatically lowering operational costs.