HugoAI is an AI research lab working on making frontier-class general intelligence radically more efficient, with the long-term goal of running AGI-class systems directly on everyday personal devices.
Capability is usually bought with scale. We think a large share of it can be bought with engineering of the whole system: the model, the runtime, the agent harness, deterministic components, and verification, measured together on one ordinary machine.
Most of the time an agent spends is not "thinking". It is re-reading context, calling tools and recovering from mistakes. Each of these is an engineering problem we can measure and remove.
When a step can be done deterministically and verified, it should not cost a single generated token. The model plans; a verified core executes.
Every number on this page comes from one laptop: MacBook Pro, M4 Pro, 24 GB of memory. No cloud, no data-center GPU.
Nown is a local agent for a computer: a 27-billion-parameter open-weight model running on the laptop, a harness that lets it act on files, apps and the web, and a verified core that executes and checks the steps it can. It is a research system, not a finished product.
Real time, one take, model running on the laptop. Typed requests in the Nown bar: file 20 invoices (20 s), produce a VAT summary exact to the cent (11 s), fill a supplier quote form without sending it (3 s). The supplier and invoices are fictitious test data.
Supplier quote form, 14 fields: 35–38 s through model tool calls → 2.0–2.6 s through the verified core, zero model calls.
VAT summary over 20 invoices: from a failed run to exact-to-the-cent results; 29/30 passes on real invoices (text, XML and OCR).
Median task time with the model's reasoning mode off, at identical success (22/28 vs 22/28), with 61 % fewer generated tokens.
| Experiment | Before | After | Status |
|---|---|---|---|
| File 20 invoices into a new folder and open it demo task, hidden judge | 84 s (n=1) | 19.9–21.1 s (n=6, all pass) | MEASURED |
| File invoices from the user's Downloads personal folders, 0 questions | 33–35 s | 26.4–27.4 s (n=3) | MEASURED |
| Supplier quote form, 14 fields, never submitted | 35–38 s | 2.0–2.6 s | MEASURED |
| VAT summary, 20 invoices, exact to the cent | 631 s, failed | 10–14 s (49/51 reports pass) | MEASURED |
| Eight end-to-end desktop tasks through the real app | 7/8 to 8/8 across runs | MEASURED | |
| Development missions (14 tasks × 2) | 22/28 | MEASURED | |
| Resumed-session first token, with our prefix-cache package | 15.5 s | 1.1 s (replayed requests, 30/30 units) | EXPERIMENTAL |
| Decode speed of the 27B model on the laptop | ≈ 9.8 tokens/s at 0–4k context · 46 % of memory bandwidth used | MEASURED | |
| Decode speed target on the same laptop | 16 tokens/s | RESEARCH GOAL | |
Machine: MacBook Pro M4 Pro, 24 GB, macOS. Model: 27B dense, UD-Q3_K_XL (13.1 GB). Dates: 21–29 September 2026. Methodology and raw tables on GitHub.
Runs where sending a form was requested: the system asked first, every time.
Runs against a booby-trapped web page: nothing sent, the trap never obeyed.
Deletion attacks on personal folders that succeeded, before and after our safety work.
Outbound non-local network connections across 663 samples during local operation.
Mutants killed by the test suite (523 tests): deliberately broken code must be caught.
The next capabilities are being built with the same method: nothing is announced before it passes the same measurement gates.
Nown operating any desktop application like a person. First proof: a text-editing task in a native app went from 631 s (failed) to 84 s (passed) once the core took over the right steps.
Multi-file changes on real codebases: plan first, then targeted edits in fresh short contexts, checked by the project's own tests.
Today the 27B model uses 46 % of the laptop's memory bandwidth. Our target is 16 tokens/s on the same machine.
Every change is packaged and fingerprinted before it runs; a measurement refers to an exact code tree, never to "the latest version".
Each change must pass fixed gates: task judges the agent cannot see, safety scenarios, and mutation testing, before it is installed.
We measure end-to-end time on the target laptop, including the model, the tools and the verification, and we attribute gains to the component that produced them.
Nown: 27B local model, verified core, measured safety. Proven on office tasks.
Reliable computer use and coding, faster decode on the same hardware, fewer model calls per task.
New weight representations, fast learned shortcuts for routine steps, smaller models where they hold up.
Frontier-level general intelligence on a phone, offline.
Hugo Sulfour, founder and technical lead. Designed and built Nown end to end: local inference integration, the agent harness, the verified core, the safety layer, and the measurement pipeline behind every number on this page. HugoAI is raising capital to accelerate this research direction.