Cascadity
Demo content: every story, outlet, person and ticker here is made up for tuning the site.
← AI
82

Lumen Labs open-sources a 9B model that runs on a phone at conversational speed

The benchmark table is less interesting than the quantization notes buried in section 4.

Lumen Labs released weights for Lumen-9, a 9-billion-parameter model it says holds 28 tokens/sec on a two-year-old flagship phone. License is permissive for companies under 50M users. Independent evals so far land it between last year's mid-size models on reasoning, weaker on long context.

Read it at the source
Cascadic Analysis 3
Product SparkOffline-first assistants just became a weekend project

What product does this inspire or accelerate?

A phone-resident model at this speed makes private, no-signal assistants viable: field service, trail guides, clinical note-taking in basements. Expect a wave of "works in airplane mode" apps.

Confidence: High
3
Investment ImpactPressure on per-token inference pricing

Where to jump in, or out? (Not financial advice.)

If good-enough runs locally, the low end of paid APIs gets squeezed. Watch mid-tier inference resellers; edge-chip designers are the likely beneficiaries.

Confidence: Medium
31
Wade InTry it on your own phone tonight

How can you try this yourself this weekend?

The reference app is in the repo. Budget 5.1 GB of storage and turn off low-power mode.

18
Rabbit Holes