Summary

Israeli AGI research company Sapient Intelligence has launched HRM-Text, an open-source 1-billion-parameter reasoning language model that takes a fundamentally different approach to AI architecture. Instead of the standard Transformer’s single forward pass, HRM-Text uses a hierarchical recurrent architecture with two nested stacks that perform multiple reasoning steps in continuous latent space before producing any output.

The key claim is efficiency: HRM-Text was trained on just 40 billion tokens of structured data — up to 1,000x fewer than the 4 to 36 trillion tokens typically required by competing models. The entire model can be trained in approximately one day for around $1,000, compared to the hundreds of millions spent training frontier LLMs. Despite its compact 0.6 GiB footprint (small enough to run on a smartphone), HRM-Text delivers competitive benchmark results: 56.2% on MATH, 81.9% on ARC-Challenge, 82.2% on DROP, and 60.7% on MMLU.

The model separates reasoning from language generation — mirroring how the human brain approaches problem-solving — and adjusts the depth of its internal reasoning based on task complexity without increasing model size.

Sources

Commentary

Bold claims deserve healthy skepticism, but HRM-Text is interesting for the right reasons. The AI industry has been locked in a scaling arms race — bigger models, more data, more GPUs — and HRM-Text asks whether that’s the only path forward. A model that fits on a phone and costs $1,000 to train, while hitting respectable (if not frontier) benchmarks, is a meaningful counterpoint to the trillion-dollar infrastructure buildout happening at OpenAI, Google, and Meta.

The architectural innovation matters more than the benchmarks. Separating reasoning from language generation and performing iterative computation in latent space before generating tokens is conceptually closer to how deliberate thinking works. Whether this approach scales to frontier-level performance remains to be seen, but even if HRM-Text stays in its lane as an efficient edge model, it opens the door for AI deployment in contexts where cloud inference is impractical or undesirable — offline devices, privacy-sensitive applications, and resource-constrained environments.

By Allan