Pokee AI Releases Pokee-Isaac 28B v0 With 10-Million Token Context on a Single RTX 4090


Pokee AI released Pokee-Isaac 28B v0 on August 3, 2026—an agentic large language model designed to handle massive datasets locally. The model is notable for claiming a 10-million-token context window while remaining deployable within customer boundaries on consumer-grade hardware, specifically a single NVIDIA RTX 4090 GPU.
By supporting a 10M-token capacity, Pokee-Isaac 28B is built to digest entire software repositories, multi-year email chains, or massive corporate documentation sets in a single pass, bypassing the need for complex Retrieval-Augmented Generation (RAG) pipelines. According to Pokee AI's technical report, the model achieved a 93.3% performance score on the RULER benchmark at its maximum 10-million-token limit, showing high retrieval accuracy across long sequences, alongside strong results on the Berkeley Function Calling Benchmark (BFCL v4).
While the technical results are promising, some industry observers have pointed out that Pokee AI has not fully disclosed the specifics of its memory compression and attention mechanisms. However, the ability to run such long-context models privately on-device represents a significant development for industries like finance and healthcare that cannot utilize cloud-based APIs due to strict data compliance. The launch adds momentum to the local, privacy-first agent trend seen in other recent releases like Liquid AI's LFM2.5 and inclusionAI's Ling 3.0 Flash.


