HugoAI
AI research lab

Frontier-class intelligence, radically more efficient.

HugoAI is an AI research lab working on making frontier-class general intelligence radically more efficient, with the long-term goal of running AGI-class systems directly on everyday personal devices.

RESEARCH GOAL AGI in your pocket — general intelligence at the level of the best frontier models, first on an ordinary computer, then on a smartphone, working even offline. This is where we are heading, not something we claim to have reached.

The thesis

Capability is usually bought with scale. We think a large share of it can be bought with engineering of the whole system: the model, the runtime, the agent harness, deterministic components, and verification, measured together on one ordinary machine.

Efficiency of the system, not just the model

Most of the time an agent spends is not "thinking". It is re-reading context, calling tools and recovering from mistakes. Each of these is an engineering problem we can measure and remove.

Let the model do only what needs a model

When a step can be done deterministically and verified, it should not cost a single generated token. The model plans; a verified core executes.

Measured on hardware people already own

Every number on this page comes from one laptop: MacBook Pro, M4 Pro, 24 GB of memory. No cloud, no data-center GPU.

Nown: our first experimental system

Nown is a local agent for a computer: a 27-billion-parameter open-weight model running on the laptop, a harness that lets it act on files, apps and the web, and a verified core that executes and checks the steps it can. It is a research system, not a finished product.

Real time, one take, model running on the laptop. Typed requests in the Nown bar: file 20 invoices (20 s), produce a VAT summary exact to the cent (11 s), fill a supplier quote form without sending it (3 s). The supplier and invoices are fictitious test data.

Model27B dense open-weight model, 3-bit quantized (13 GB), local inference.
Runtimellama.cpp with our own prefix-cache work to avoid re-reading context.
HarnessAgent loop with tools for files, code, apps (Computer Use) and the browser.
Verified coreDeterministic actions and on-disk verification; the model is called only where it adds value.
Safety layerAccess only to what the user's own request names; content from pages or files can never grant access.

What we have measured

MEASURED logged runs, replicated or n ≥ 2 EXPERIMENTAL measured, but preliminary or not yet live RESEARCH GOAL not measured, where we are heading
MEASURED
≈ ×15

Supplier quote form, 14 fields: 35–38 s through model tool calls → 2.0–2.6 s through the verified core, zero model calls.

MEASURED
631 s → 11 s

VAT summary over 20 invoices: from a failed run to exact-to-the-cent results; 29/30 passes on real invoices (text, XML and OCR).

MEASURED
−48 %

Median task time with the model's reasoning mode off, at identical success (22/28 vs 22/28), with 61 % fewer generated tokens.

ExperimentBeforeAfterStatus
File 20 invoices into a new folder and open it
demo task, hidden judge
84 s (n=1)19.9–21.1 s (n=6, all pass)MEASURED
File invoices from the user's Downloads
personal folders, 0 questions
33–35 s26.4–27.4 s (n=3)MEASURED
Supplier quote form, 14 fields, never submitted35–38 s2.0–2.6 sMEASURED
VAT summary, 20 invoices, exact to the cent631 s, failed10–14 s (49/51 reports pass)MEASURED
Eight end-to-end desktop tasks through the real app7/8 to 8/8 across runsMEASURED
Development missions (14 tasks × 2)22/28MEASURED
Resumed-session first token, with our prefix-cache package15.5 s1.1 s (replayed requests, 30/30 units)EXPERIMENTAL
Decode speed of the 27B model on the laptop≈ 9.8 tokens/s at 0–4k context · 46 % of memory bandwidth usedMEASURED
Decode speed target on the same laptop16 tokens/sRESEARCH GOAL

Machine: MacBook Pro M4 Pro, 24 GB, macOS. Model: 27B dense, UD-Q3_K_XL (13.1 GB). Dates: 21–29 September 2026. Methodology and raw tables on GitHub.

Safety and verification, measured

MEASURED
26/26

Runs where sending a form was requested: the system asked first, every time.

MEASURED
14/14

Runs against a booby-trapped web page: nothing sent, the trap never obeyed.

MEASURED
17/20 → 0/20

Deletion attacks on personal folders that succeeded, before and after our safety work.

MEASURED
0

Outbound non-local network connections across 663 samples during local operation.

MEASURED
362/362

Mutants killed by the test suite (523 tests): deliberately broken code must be caught.

What's next, already in the lab

The next capabilities are being built with the same method: nothing is announced before it passes the same measurement gates.

IN DEVELOPMENT

Universal Computer Use

Nown operating any desktop application like a person. First proof: a text-editing task in a native app went from 631 s (failed) to 84 s (passed) once the core took over the right steps.

IN DEVELOPMENT

Coding on real projects

Multi-file changes on real codebases: plan first, then targeted edits in fresh short contexts, checked by the project's own tests.

RESEARCH GOAL

Twice the speed, same laptop

Today the 27B model uses 46 % of the laptop's memory bandwidth. Our target is 16 tokens/s on the same machine.

Method

Frozen before measured

Every change is packaged and fingerprinted before it runs; a measurement refers to an exact code tree, never to "the latest version".

Gates, not impressions

Each change must pass fixed gates: task judges the agent cannot see, safety scenarios, and mutation testing, before it is installed.

One machine, whole system

We measure end-to-end time on the target laptop, including the model, the tools and the verification, and we attribute gains to the component that produced them.

Research roadmap

NOW

A capable agent on a laptop

Nown: 27B local model, verified core, measured safety. Proven on office tasks.

NEXT

Whole-system efficiency

Reliable computer use and coding, faster decode on the same hardware, fewer model calls per task.

THEN

Less memory per unit of intelligence

New weight representations, fast learned shortcuts for routine steps, smaller models where they hold up.

GOAL

AGI in your pocket

Frontier-level general intelligence on a phone, offline.

Founder

Hugo Sulfour, founder and technical lead. Designed and built Nown end to end: local inference integration, the agent harness, the verified core, the safety layer, and the measurement pipeline behind every number on this page. HugoAI is raising capital to accelerate this research direction.