Read the memory indicator
The toolbar shows Noema’s current app memory usage alongside the Mac’s total physical RAM. Hover over the indicator for an explanation.
The second value is total installed RAM, not currently free memory. The readout measures the app as a whole, including models, context, documents, and other active work.
Model file size measures storage. It is not a direct measure of runtime RAM use.
Run a benchmark
- Save the model settings you want to measure.
- Open Tools → Benchmarking Center.
- Select an installed, supported model.
- Choose Run Benchmark.
- Review the result and measurement time.
Noema uses the selected model’s saved settings and reloads the model when needed to match them.
Understand the results
| Metric | Meaning |
|---|---|
| Generation | Output-token throughput after generation begins. Higher is faster. |
| Prefill | Prompt-processing throughput. |
| First token | Time until the first output arrives. |
| Total time | Duration of the measured request. |
| Peak memory | Highest app memory usage observed during the run. |
| Added memory | Increase above the benchmark’s starting memory usage. |
Depending on the runtime, token rates may use native timings or estimates.
Compare results fairly
The benchmark uses a short request with a 512-token output limit. Everyday chats can take longer because they include conversation history, retrieved passages, tools, workspace preparation, or longer reasoning.
Generation speed excludes prompt-processing time. For comparisons, keep the model, quantization, settings, device conditions, and competing workloads as consistent as possible.