← All modelsOpen weights / Model 02Available / August 12, 2026

Noema model family

Noema 1.5 2B

More capable. More reliable. Still completely local. An open 2B model with recovered knowledge, stronger code, precise instructions, and shorter reasoning.

Why 1.5 exists

A better balance, carried entirely in the weights.

The previous Noema release developed useful strengths in mathematics, code, and structured responses, but paid a measurable price in broad knowledge. Noema 1.5 was trained specifically to recover that deficit without giving up the capabilities that made Noema useful. The result is a compact model with better knowledge retention, stronger code generation, more precise instruction following, improved multi-turn consistency, and shorter reasoning traces. No retrieval, external tools, or inference-time answer repair were used in the reported results; the improvements live in the weights.

  • 01Recovered broad knowledge
  • 02Precise instruction following
  • 03Multi-turn constraint retention
  • 04Efficient local reasoning
Status
Available now
Released
August 12, 2026
Parameters
Approximately 2B
Starting checkpoint
Noema 2B
Architecture
Hybrid: Gated DeltaNet + attention
Layers / hidden
24 / 2,048
Context
262K native · 24,576 independently evaluated
Training
Restoration-aware on-policy distillation
Modes
Non-thinking · thinking
Primary language
English
License
Apache 2.0
Noema model library showing downloadable local models on Mac
Inside Noema

Download Noema 1.5 into the Noema model library, choose a device-aware runtime preset, and keep every prompt and response on hardware you control.

The new balance

Useful behavior moved forward.

Matched development gates show progress from Noema 2B; the final lockbox compares the frozen release with stock Qwen3.5-2B.

+2.86 pts

Knowledge recovered

53.00 vs 50.14 on the matched 700-question development set, alongside a 1.42-point gain on the pooled knowledge composite.

MMLU-Pro dev
+7.02 pts

Instructions followed

72.09 vs 65.06 for stock Qwen3.5-2B on strict prompt-level instruction following; the paired result is statistically significant.

IFEval final
+3.66 pts

Code that passes

Pass@1 increased from 50.00 to 53.66 over Noema 2B and remained favorable against the stock foundation in final testing.

HumanEval+ dev
+21.1%

Complete conversations

Relative increase in conversations satisfying every tested turn: 17.88% for Noema 1.5 vs 14.77% for stock Qwen3.5-2B.

Multi-IF final
−13.5%

Shorter reasoning

Mean completion length fell from 7,969 to 6,895 tokens while the thinking-accuracy point estimate remained favorable.

Thinking tokens
+14 pts

Verified mathematics

The verified mathematics development gate moved from 78.00 to 92.00 over the previous Noema release.

Math dev

Two frozen comparison stages

Benchmarks

Percentages unless noted · identical settings within each comparison

+7.02percentage points

Strict instructions

Final IFEval prompt-strict improvement over stock Qwen3.5-2B; statistically significant at p = 0.000475.

+21.1%relative success

Every turn satisfied

Relative lift in complete three-turn Multi-IF conversations over the stock model.

−13.5%completion tokens

Less reasoning overhead

Shorter mean thinking traces with a favorable accuracy point estimate in the final evaluation.

Matched development gates

Versus Noema 2B

06 measuresHigher is better

Pooled knowledge

+1.42 pp
composite · 1,900 questions

MMLU-Pro

+2.86 pp
development · 700 questions

HumanEval+

+3.66 pp
pass@1

IFEval

+2.77 pp
prompt strict

Verified math

+14.00 pp
development

Novel constraints

+5.83 pp
prompt strict

Paired development and selection results against the exact Noema 2B starting checkpoint. These gates informed model selection; they are not a third arm in the untouched final evaluation.

One-time final evaluation

Versus stock Qwen3.5-2B

09 measuresHigher is better · lower is better for tokens

IFEval

+7.02 pp
prompt strict · 541 prompts

Multi-IF

+2.36 pp
mean per turn · 4,501 conversations

Multi-IF

+3.11 pp
all three turns · 4,501 conversations

HumanEval+

+3.05 pp
pass@1 · 164 problems

Thinking accuracy

+1.33 pp
native · 300 prompts

Thinking tokens

−13.5%
mean completion · 300 prompts

MMLU-Pro

−0.71 pp
knowledge retention · 11,332 questions

GPQA-Diamond

−5.05 pp
198 questions

IFBench

−3.67 pp
prompt loose · 300 prompts

The frozen Noema 1.5 candidate and stock Qwen3.5-2B used identical prompts, generation settings, and graders. MMLU-Pro was the preregistered knowledge-retention endpoint; point estimates are reported even when they were not statistically resolved.

Publisher-reported model context

Competitive below 2B. Directionally close to selected larger models.

Noema 1.5 occupies a strong middle ground: balanced knowledge and science reasoning, strict instruction following, and a compact package designed for local deployment.

Protocol noteDirectional context, not a leaderboard. External scores are reported by each model publisher and were not produced in Noema's controlled harness. Prompt templates, reasoning modes, sampling, token budgets, benchmark revisions, quantization, and graders may differ. Only stock Qwen3.5-2B was evaluated under Noema's paired protocol above.

The open sub-2B field

A highly competitive, unusually balanced compact model.

Published results from current openly available models below two billion total parameters. The protocol label beside each model is part of the comparison.

Noema 1.5 2BFinal L8 · prompt-strict IFEval
Noema final evaluation
LFM2.5-1.2B ThinkingThinking · averaged IFEval
Liquid AI model card
EXAONE 4.0 1.2BNon-reasoning
LG AI model card
EXAONE 4.0 1.2BReasoning
LG AI model card

MMLU-Pro

Knowledge and reasoning
Higher is better

GPQA-Diamond

Graduate-level science
Higher is better

IFEval

Instruction following
Higher is better

ReadoutNoema reports stronger MMLU-Pro and GPQA-Diamond results than LFM2.5-1.2B Thinking and narrowly exceeds EXAONE's non-reasoning results. EXAONE leads both knowledge measures in reasoning mode, while LFM leads the instruction-focused result.

Selected larger edge models

Context beyond the two-billion-parameter line.

Selected publisher results from larger local and edge-oriented models, shown on the same score scale for easier comparison.

Noema 1.5 2BFinal L8 · prompt-strict IFEval
Noema final evaluation
Phi-4 Mini0-shot CoT
Microsoft model card
Gemma 4 E2BPublisher evaluation
Google model card
Ministral 3 3BReasoning
Mistral model card
Nemotron 3 Nano 4BFP8 · reasoning on
NVIDIA model card
Qwen3.5-4BPublisher evaluation
Qwen model card

MMLU-Pro

Knowledge and reasoning
Higher is better

GPQA-Diamond

Graduate-level science
Higher is better

IFEval

Instruction following
Higher is better

ReadoutNoema's reported MMLU-Pro result slightly exceeds Phi-4 Mini's. Gemma, Ministral, Nemotron, and Qwen retain advantages on their published measures.

Evaluation discipline

Useful gains, reported with their edges intact.

The final comparison was run once after the candidate was frozen. No retrieval, external tools, answer repair, or outside model calls were used to produce benchmark answers.

Strict instructions
p = 0.000475The +7.02-point IFEval gain was statistically significant.
Complete multi-turn success
p = 4.44e−8The +3.11-point all-three-turn gain was statistically significant.
Knowledge retained
11,332 questionsMMLU-Pro passed the preregistered non-inferiority gate.

Evaluation controls

  • Frozen model and tokenizer revisions
  • Identical prompts and settings across both arms
  • BF16 inference on NVIDIA L40S hardware
  • Pinned official graders and paired statistics
  • Deterministic keys and exact-count verification

Measured limits

  • GPQA-Diamond and IFBench finished below stock in the final point estimates.
  • HumanEval+ and thinking accuracy were favorable but not statistically resolved.
  • Evaluation was primarily English and text-only; context was independently validated to 24,576 tokens.

Apache 2.0 / Open weights

More disciplined AI, on the device you already own.

Open weights under Apache 2.0. Download the official model and run it privately in Noema or another compatible local runtime.

Download Noema 1.5 2B