Pokee-Isaac 28B 10M-Token Context AI Stays Inside Your Boundary

Long‑horizon agents keep every tool output, observation and reasoning step in their context window. As the window grows, the model must hold massive amounts of information and stay coherent across it. Cloud‑based large models are the only option that currently offers both a huge context window and the compute to keep it usable, but regulated industries, public‑sector agencies and on‑device applications cannot send data outside their boundaries. This forces teams to either truncate context, lose fidelity, or abandon ambitious agentic workflows.

Pokee‑Isaac 28B solves the problem by delivering a 28‑billion‑parameter text‑only foundation model with a 10‑million‑token context window that is licensed to run inside a VPC, on‑premises cluster or directly on edge hardware. Independent measurements show 93.3 % on the RULER benchmark at the full 10 M‑token length, matching the best cost‑optimized cloud baselines while fitting on a single GPU. Prefill throughput scales with context—rising from ~42 k tokens/s at 1 M tokens to ~137 k tokens/s at 10 M tokens—so a ten‑fold longer prompt costs only about three times the time‑to‑first‑token. Decode speed stays flat near 335 tokens/s, making output generation predictable regardless of resident context.

Because the model weights are not open‑weight, deployment is through a licensed API that can be placed behind any organizational firewall. This fits mid‑size and enterprise teams that already operate their own inference stacks, device OEMs, and any organization subject to data‑ residency rules—healthcare, finance, defense, legal, pharma and semiconductor R&D. Use cases include whole‑repository code review, multi‑year contract analysis, incident forensics over full log archives, and long‑running tool agents that never need summarization or pruning.

In short, Pokee‑Isaac 28B gives regulated and edge‑bounded organizations the ability to keep massive context locally, retain high agentic performance, and avoid costly data egress—all without sacrificing speed or security.

#AI #LLM #OnPrem #EdgeAI #ContextWindow #SecureAI