# Noema Documentation --- --- > Complete Markdown documentation export. Individual pages are available at `/docs/.md`. --- --- --- title: "Noema Constellation" description: "Sync Noema chats and projects between your Apple devices, then use models installed on an awake Mac from iPhone, iPad, or Vision Pro." version: "Noema 3.6+" platforms: ["iPhone", "iPad", "Mac", "Vision Pro"] reviewed: "July 26, 2026" canonical: "https://noemaai.com/docs/constellation" --- # Noema Constellation Your chats follow you, and your Mac’s models become available across your personal devices. ## One feature, two kinds of continuity Constellation brings your Noema devices together. Chats and projects sync through your private iCloud database, while models installed on an awake Mac can be selected from another device without copying the model weights. There is no pairing code, host ID, or IP address to type. Devices signed into the same Apple Account discover one another automatically. ## Set up your Mac 1. Open Noema on the Mac that stores the models you want to use. 2. Open **Settings → Sync**. 3. Turn on **Sync chats** if you want chat and project continuity. 4. Turn on **Reach this Mac from anywhere** to allow remote model access. 5. Choose when Noema may keep the Mac awake. 6. Leave Noema running and keep the Mac awake. > Remote Access is optional and remains off until you enable it. Noema explains the network behavior before it turns on. ## Connect from iPhone, iPad, or Vision Pro 1. Make sure the device uses the same Apple Account as the Mac. 2. Open **Stored → Constellation**. 3. Turn on **Sync chats** and **Use your Mac’s models**. 4. Wait for the Mac to show as ready. 5. Find a model under the Mac and press **Use**. Noema starts loading the model and chooses a connection route before the first message. Chat opens immediately; if the model is large, the first request waits for that in-progress load to finish. ## What syncs - Chats, messages, titles, favorites, and bookmarks. - Chat instructions, scratchpads, modes, answer styles, and tool permissions. - Projects, project instructions, source references, defaults, and archive state. - Tool history, web results, and route information shown in a transcript. - Deletions and regenerated turns. ## What stays on each device - Model weights and downloads. - Datasets, PDFs, source files, indexes, and embeddings. - Per-device MCP server selections and runtime settings. - Photo and media attachment files. Project source references can sync, but the corresponding local source must also exist on a device before that device can retrieve from it. ## Force an immediate sync Open the chat drawer on iPhone or iPad and pull down. Noema pushes local changes, fetches remote changes, merges them, and sends any merge result before the refresh indicator finishes. > Off-Grid Mode pauses Constellation sync and remote access. ## Related documentation - [Constellation Models, Tools & Photos](https://noemaai.com/docs/constellation-models-and-tools) - [Constellation Connections & Privacy](https://noemaai.com/docs/constellation-connections) - [Constellation Model Ejection & Help](https://noemaai.com/docs/constellation-troubleshooting) - [Privacy & Network Activity](https://noemaai.com/docs/privacy-and-network-activity) --- --- title: "About Noema" description: "Learn how Noema keeps local AI private while making optional network features explicit and controllable." version: "Noema 3.6+" platforms: ["iPhone", "iPad", "Mac", "Vision Pro"] reviewed: "July 26, 2026" canonical: "https://noemaai.com/docs/about-noema" --- # About Noema What local-first means, what can stay offline, and when Noema uses a network connection. ## Local-first, with the route in your control Noema is a native AI workspace for Apple devices. Local models, conversations, document indexes, and many tools can stay on your hardware and continue working without an internet connection. Some features intentionally use a network: Web Search, model and dataset downloads, remote endpoints, Constellation, remote Autopilot routing, enterprise services, and configured speech or tool providers. Noema shows and controls those paths instead of pretending they do not exist. Constellation can use an awake Mac’s models from another personal device. Noema prefers the local network, a configured Private Link, or an encrypted direct path; Noema Bridge and private CloudKit exchange are fallbacks when those routes are unavailable. > For a feature-by-feature view of what stays local and what may be sent, read Privacy & Network Activity. ## What works offline - Chats with an installed GGUF, MLX, ExecuTorch, Core ML, or supported on-device Apple Foundation Model. - Retrieval over downloaded documents, datasets, and Knowledge Packs after an embedding model has indexed them. - Local tools including Python, Memory, Calculator, Unit Converter, Charts, and supported device integrations. - Local Whisper transcription and optional local neural voice after their models are installed. ## What requires a deliberate connection - Searching or reading public web sources. - Sending a turn to a remote endpoint or stronger Autopilot model. - Using Constellation to sync chats or reach an awake Mac model. - Downloading models, datasets, voices, embeddings, or Knowledge Packs. - Using a remote speech engine, enterprise workspace, or external tool server. ## Off-grid Mode Turn on Settings → Off-grid Mode when you want Noema to block external HTTP and HTTPS requests made through its network stack. Local generation, retrieval, and local tools remain available; network-backed features stop or fall back locally when possible. Off-grid Mode does not sandbox separate software or operating-system services. Review the privacy page for those boundaries. ## Related documentation - [Noema Constellation](https://noemaai.com/docs/constellation) - [Quick Start](https://noemaai.com/docs/quick-start) - [Privacy & Network Activity](https://noemaai.com/docs/privacy-and-network-activity) --- --- title: "Quick Start" description: "Set up Noema, download or activate your first model, and begin a private local chat in a few minutes." version: "Noema 3.6+" platforms: ["iPhone", "iPad", "Mac", "Vision Pro"] reviewed: "July 26, 2026" canonical: "https://noemaai.com/docs/quick-start" --- # Quick Start Install Noema, choose a starter model, and send your first fully local message. ## What you need A chat model is required to generate replies. An embedding model is separate and only required when you want to search documents, datasets, or Knowledge Packs. - iPhone or iPad with iOS or iPadOS 18 or later and an A12 Bionic chip or newer. - Apple-silicon Mac with macOS 26 or later. - Apple Vision Pro with visionOS 26 or later. - Enough free storage for the model you choose; sizes range from hundreds of megabytes to many gigabytes. - Wi-Fi for the initial model download is recommended. ## Start your first local chat 1. Install Noema from the App Store and open it. 2. Choose the guided welcome flow or open the full setup controls. 3. Download the recommended Qwen 3.5 2B GGUF model, or choose another compatible model in Explore. 4. Wait until the model is marked ready in Stored. 5. Select it in Chat and send your first message. > Qwen 3.5 2B is a balanced starting point. More memory can support larger models; smaller models and lighter quantizations usually load faster and use less energy. ## Optional document memory Install the embedding model when you want Noema to index and retrieve your own material. - Chat with PDFs, EPUBs, Markdown, text, or JSON files. - Search downloaded datasets and Knowledge Packs. - Retrieve source passages for citation-aware answers. ## Use a model already on your Mac If you already keep larger models on a Mac, you can use **Stored → Constellation** instead of downloading another copy to your phone. Turn on Remote Access on the Mac, select its model, and Noema will load it and choose the best available connection. ## Stay fully local After models and datasets are downloaded, local chat and retrieval can work offline. Off-grid Mode blocks external HTTP and HTTPS traffic initiated through Noema. Web Search, remote endpoints, Constellation sync and remote access, remote Autopilot routing, new downloads, and remote speech services require network access and will not operate normally in Off-grid Mode. ## If a model will not load - Close memory-heavy apps. - Choose a smaller model or lighter quantization. - Reduce context length or choose the Battery Saver preset. - Follow the RAM guidance shown on the model page. - Open Settings → Diagnostics & Tools for device, storage, model, and runtime checks. ## Related documentation - [Installation](https://noemaai.com/docs/installation) - [Noema Constellation](https://noemaai.com/docs/constellation) - [Running Models Locally](https://noemaai.com/docs/running-llms) --- --- title: "Installation" description: "Install Noema on supported Apple devices and prepare the right amount of storage and memory for local models." version: "Noema 3.5+" platforms: ["iPhone", "iPad", "Mac", "Vision Pro"] reviewed: "July 20, 2026" canonical: "https://noemaai.com/docs/installation" --- # Installation Current system requirements, App Store installation, and first-model setup. ## System requirements - iPhone or iPad: iOS or iPadOS 18 or later is required; supported hardware starts at A12 Bionic. - Mac: Apple silicon with macOS 26 or later. - Apple Vision Pro: visionOS 26 or later. - Storage: keep enough headroom for the selected model, its context, and temporary download data. Five gigabytes is useful recommended headroom, not a universal fixed requirement. ## Install from the App Store 1. Open Noema’s App Store listing on the device you want to use. 2. Tap Get and authenticate with Face ID, Touch ID, or your Apple Account password. 3. Open Noema and complete the welcome flow. 4. Choose the guided setup or open Explore to select a model yourself. ## Recommended starting model Start with Qwen 3.5 2B GGUF. It is the current guided-setup model and balances response quality, load time, memory use, and compatibility across supported phones and tablets. The embedding model is optional for ordinary chat. Install it only when you want document retrieval, downloaded datasets, or Knowledge Packs. ## Plan storage and memory - Download large assets over Wi-Fi when possible. - Do not assume parameter count alone determines fit; quantization, context length, runtime, and multimodal dependencies also consume memory. - Keep free storage available for resumable download data and indexing work. - Use the model page’s fit guidance before loading a larger model. ## Related documentation - [Quick Start](https://noemaai.com/docs/quick-start) - [Downloading Models](https://noemaai.com/docs/downloading-models) - [Support](https://noemaai.com/docs/support) --- --- title: "Running Models Locally" description: "Choose among Noema’s local runtimes using memory fit, quantization, context, and capability guidance." version: "Noema 3.5+" platforms: ["iPhone", "iPad", "Mac", "Vision Pro"] reviewed: "July 20, 2026" canonical: "https://noemaai.com/docs/running-llms" --- # Running Models Locally How GGUF, MLX, ExecuTorch, Core ML, and Apple Foundation Models differ in Noema 3.5. ## Local runtime choices | Runtime | Best fit | What to know | | --- | --- | --- | | GGUF | Broad compatibility and detailed tuning | Flexible quantizations, context controls, projectors, and advanced runtime options. | | MLX | Apple-silicon optimization | Efficient Apple-native execution with a more opinionated settings surface. | | ExecuTorch | Packaged mobile deployments | Compatibility and available controls depend on the exported model bundle. | | CML / Core ML | Apple neural-engine and packaged model workflows | Dependencies and supported capabilities are model-specific. | | AFM | Apple’s built-in Foundation Model | Availability depends on device and OS; it is activated, not downloaded like a model file. | ## Choose by fit, not parameter count alone - RAM fit includes model weights, key-value cache, context length, runtime overhead, and vision or audio dependencies. - A lighter quantization reduces memory use but can change quality and speed. - Long context can consume more memory than expected even with a small model. - Vision, tool calling, reasoning, and audio support are separate capabilities; inspect badges and model notes before downloading. ## Load and run a model 1. Open Explore and choose a compatible model and quantization. 2. Review license, provenance, capability badges, dependencies, and estimated fit. 3. Download the model and any required projector or companion files. 4. Open Stored, review its runtime settings, and load it. 5. Select it in Chat and monitor the context gauge and runtime receipt. ## Performance and energy Use a runtime preset instead of enabling system Low Power Mode as a model-tuning strategy. Battery Saver reduces the workload through Noema’s own settings; Balanced is a good default; Max Speed favors throughput when heat and power use are acceptable. - Reduce context length before assuming the model is too large. - Close memory-heavy apps and avoid loading multiple large local models unless the fit estimate allows it. - Benchmark on the device where the model will actually run. ## Related documentation - [Downloading Models](https://noemaai.com/docs/downloading-models) - [Model Settings](https://noemaai.com/docs/model-settings) - [Quick Start](https://noemaai.com/docs/quick-start) --- --- title: "Downloading Models" description: "Find compatible local models, inspect capabilities and provenance, and manage resumable downloads in Noema." version: "Noema 3.5+" platforms: ["iPhone", "iPad", "Mac", "Vision Pro"] reviewed: "July 20, 2026" canonical: "https://noemaai.com/docs/downloading-models" --- # Downloading Models Formats, fit guidance, dependencies, licenses, and the path from Explore to Stored. ## What can be downloaded Noema 3.5 discovers and downloads compatible GGUF, MLX, ExecuTorch, and Core ML model packages. Apple Foundation Models are supplied by the operating system and appear when supported; they are not normal model downloads. ## Read the model page before downloading - Compatibility and estimated RAM fit for the current device. - Quantization, file size, architecture, and context guidance. - Vision, tool-calling, reasoning, or audio capability badges. - Required projector, tokenizer, or other companion dependencies. - Publisher, source repository, license, and provenance notes. ## From Explore to Stored 1. Open Explore → Models and search or browse the catalog. 2. Choose a model, then select a compatible file or quantization. 3. Review fit guidance and dependencies before tapping Download. 4. Monitor progress. Interrupted supported downloads can resume rather than restart from zero. 5. Open Stored when the model is ready, adjust its settings, and load it. ## Troubleshooting downloads - Free enough storage for the final file and temporary transfer data. - Keep the source license and gated-repository requirements in mind. - If a vision model loads without image support, check for its required projector. - If a file is marked incompatible, choose a supported architecture or runtime instead of forcing it. - Off-grid Mode blocks new model and dependency downloads. ## Related documentation - [Running Models Locally](https://noemaai.com/docs/running-llms) - [Model Settings](https://noemaai.com/docs/model-settings) - [Installation](https://noemaai.com/docs/installation) --- --- title: "Model Settings" description: "Tune Noema models with presets, context and sampling controls, RAM guards, benchmarking, and runtime-specific options." version: "Noema 3.5+" platforms: ["iPhone", "iPad", "Mac", "Vision Pro"] reviewed: "July 20, 2026" canonical: "https://noemaai.com/docs/model-settings" --- # Model Settings A practical guide to presets, context, sampling, helper models, and format-dependent controls. ## Start with a runtime preset | Preset | Use it for | | --- | --- | | Battery Saver | Lower sustained workload and energy use. | | Balanced | Everyday chat with sensible memory and speed trade-offs. | | Max Speed | Highest practical throughput when thermal and battery cost are acceptable. | | Max Context | Longer conversations or documents when the model and device fit. | | Vision Heavy | Image prompts that need additional multimodal headroom. | | Tool Heavy | Chats expected to use many tool definitions and results. | > You can save a custom preset after tuning a model for a repeatable workload. ## Context and RAM guards Context length affects memory, speed, and how much conversation or retrieved material the model can see. Noema estimates the working set against the current device and warns when a configuration is likely to exceed safe headroom. - Reduce context before changing many advanced controls at once. - Leave space for attachments, retrieval passages, and tool schemas. - A configuration that fits on a Mac may not fit on an iPhone with the same model file. ## Sampling and generation - Temperature, top-p, top-k, min-p, and repetition controls vary by runtime. - Reasoning and maximum-output controls appear only when supported by the selected model. - Some packaged runtimes intentionally expose fewer options than GGUF. - Reset to a preset if a heavily tuned configuration produces unstable or repetitive output. ## Advanced local features The faster GGUF settings screen can expose benchmarking, speculative decoding with a compatible helper model, and multi-token prediction options. These controls are format- and model-dependent and should be enabled only when the compatibility checks pass. Vision options appear only for models that support them. ## Apple Foundation Models AFM settings depend on the operating system and supported device. Noema does not present the obsolete Default or Permissive guardrail choice. Availability, execution mode, and any Apple cloud controls are shown only when the public OS exposes them. ## Related documentation - [Running Models Locally](https://noemaai.com/docs/running-llms) - [Downloading Models](https://noemaai.com/docs/downloading-models) - [Support](https://noemaai.com/docs/support) --- --- title: "Using Noema Overfit" description: "Run compatible mixture-of-experts models with a resident core, locally paged expert weights, and device-specific performance guidance." version: "Noema 3.6+" platforms: ["iPhone", "iPad", "Mac", "Vision Pro"] reviewed: "August 24, 2026" canonical: "https://noemaai.com/docs/noema-overfit" --- # Using Noema Overfit Run compatible Qwen, Gemma, Laguna, and DeepSeek mixture-of-experts models using a resident core and locally paged expert weights, then measure whether the model is practical on your device. ## How Overfit works Noema Overfit lets compatible mixture-of-experts models run even when their complete weights cannot fit in device memory. Overfit is not model training or “overfitting.” It reorganizes a compatible GGUF model into a complete `.noema-paged` package containing: - `resident.gguf`, with the weights required throughout inference. - One or more `experts-*.bin` files containing routed expert weights. - `manifest.json`, describing the package structure, model geometry, file hashes, expert records, and integrity information. During inference, Noema keeps the resident core and a device-sized expert bank in unified memory. When the model selects an expert that is not already present, Noema reads it from local storage, verifies it, and loads it into the bank. > Overfit provides capacity, not guaranteed speed. Storage is not RAM, and a smaller model that fits normally will usually start faster, generate faster, and consume less energy. ## Download and run an Overfit model 1. Open **Explore** and select the **GGUF** catalog. 2. Open the **Noema Overfit** paged-model collection. 3. Select a complete paged package and review its storage size, resident memory, expert-bank estimate, context limit, and estimated working set. 4. Download the package. Noema treats the manifest, resident GGUF, and expert files as one model. 5. Open the model in **Stored**, then open its settings. 6. Leave **Overfit (Paged Experts)** set to **Automatic**. 7. Run the **Canary Test** to measure the model on this device. 8. Load the model and begin a chat. > Overfit packages are models, not ordinary quantization choices. Never download only `resident.gguf`; it cannot represent the complete model without the manifest and expert files. ## Create a package on Mac Noema for macOS can convert an existing compatible MoE GGUF into a complete paged package. 1. Download and register the source GGUF. 2. In **Stored**, right-click the model. 3. Choose **Create Paged Package…** 4. Wait while Noema extracts and verifies the expert data. 5. Choose **Add to Stored** when conversion finishes. The result is a sibling `.noema-paged` directory. Conversion does not retrain, prune, or replace the model’s weights. > The original GGUF remains available, so enough storage is required for both copies during conversion. ## The Canary Test The Canary Test performs a short, repeatable local run for this package and device. - Validates a sample of the package. - Measures cached and uncached storage reads. - Loads the paged runtime. - Runs a fixed 64-token completion. - Records time to first token and generation speed. - Measures latency percentiles and long stalls. - Records expert-bank hit rate, misses per token, peak memory, and thermal state. The result belongs to the exact package fingerprint, device, storage volume, native runtime contract, and app build. Available memory is evaluated separately at launch. ## Understand the Canary result - **Paged — interactive** - **Paged — slow** - **Too slow on this device** - **Use Constellation** - **Not supported for this model** > The Canary Test is recommended rather than mandatory. Performance can change with free memory, thermal conditions, storage activity, prompt length, and model settings. ## Current compatible models As of August 24, 2026, the public [Noema Overfit repository](https://huggingface.co/NoemaAI-labs/Noema-Overfit) contains these complete packages. | Model package | Approximate download | Resident core | Expert payloads | | --- | --- | --- | --- | | DeepSeek V4 Flash 0731, UD-IQ4_NL | 136.66 GB | 7.81 GB | 128.85 GB | | Gemma 4 26B-A4B QAT, UD-Q4_K_XL | 14.36 GB | 1.40 GB | 12.96 GB | | Qwen 3.6 35B-A3B, UD-Q4_K_M | 22.15 GB | 2.57 GB | 19.57 GB | | Qwen 3.5 122B-A10B, Q4_K_M | 74.27 GB | 4.00 GB | 70.26 GB | The runtime whitelist currently covers `qwen3moe`, `qwen35moe`, `gemma4`, `laguna`, and `deepseek4`. This includes the published DeepSeek V4 Flash, Qwen 3.5, Qwen 3.6, and Gemma 4 packages. Poolside's [Laguna S 2.1 GGUF](https://huggingface.co/poolside/Laguna-S-2.1-GGUF) and [Laguna XS 2.1 GGUF](https://huggingface.co/poolside/Laguna-XS-2.1-GGUF) can be converted into complete paged packages on macOS; the public Noema collection does not currently list pre-built Laguna packages. DeepSeek V4 Flash uses the dedicated `deepseek4` runtime and native contract v4. Treat it as a model-specific package, not as a conventional Qwen or Gemma package. Dense models and arbitrary GGUF architectures are not supported. ## Privacy and safety After the model is downloaded, Overfit inference can remain entirely local. - Model weights remain on local storage. - Prompts and responses remain on the device. - No cloud inference provider is required. - Expert payloads are excluded from device backups. - Matching conversation-prefix states can be saved locally for faster follow-up launches. - Using a model from an awake Mac through Constellation remains a separate, optional choice. Packages are checked for valid geometry, safe paths, file sizes, record bounds, coverage, fingerprints, and checksums. The native runtime validates the package independently at launch. Under memory pressure, Noema can reduce prefetching, discard optional queued work, or stop generation before an unsafe memory condition. ## Current limitations - Overfit is restricted to compatible MoE architectures. - Package creation is currently macOS-only. - Paged launches currently cap context at 8,192 tokens on Mac and 4,096 tokens on iPhone, iPad, and Vision Pro. - Multimodal projectors and some utility workflows do not use the paged path yet. - Low Power Mode disables expert prefetching to favor energy use. - Prompt-state restoration only works when the package and relevant runtime configuration still match. - **Force (Experimental)** may bypass performance recommendations, but it never bypasses package-integrity validation. ## Related documentation - [Downloading Models](https://noemaai.com/docs/downloading-models) - [Running Models Locally](https://noemaai.com/docs/running-llms) - [Model Settings](https://noemaai.com/docs/model-settings) --- --- title: "Constellation Models, Tools & Photos" description: "Select a model from another Noema device, apply portable model settings, use local tools and documents, and send photos to supported vision models." version: "Noema 3.6+" platforms: ["iPhone", "iPad", "Mac", "Vision Pro"] reviewed: "July 26, 2026" canonical: "https://noemaai.com/docs/constellation-models-and-tools" --- # Constellation Models, Tools & Photos The model runs on the Mac while the requesting device keeps control of tools, local knowledge, permissions, and the chat. ## Read a remote model row Models installed on another device appear in Stored under that device’s name. - Loaded or available state. - Provider and runtime. - Quantization, size, and context length. - Measured generation speed. - An eye icon for usable image input. - A tool icon for reported native tool calling. A sleeping or unreachable Mac is shown once at the device level. Its individual models are not presented as if they were currently usable. ## Configure before using Press the model row to open its settings. Sampling, context length, stop sequences, prompt caching, reasoning, and other portable inference choices can follow the session to the Mac. The Mac keeps control of hardware-specific choices such as model paths, GPU offload, thread count, and device-dependent loading policy. Press **Use** to load the model on the Mac and switch Chat to it. The route is chosen during this activation instead of waiting for the first message. ## Tools run where your data is Constellation sends the selected model the tools currently allowed by the requesting device. When the model asks to use one, that device executes it and returns the result. This design lets a model running on the Mac use a dataset or PDF indexed on the requesting device without copying the entire library to the Mac. Only context or tool results needed for the active conversation travel through the selected route. - Web Search. - Python. - Memory. - Calculator and Unit Converter. - Dataset search and retrieval. - PDF navigation. - Calendar. - Chart rendering. - Selected MCP tools. > Tool availability still depends on model compatibility, per-chat permissions, OS permissions, installed resources, workspace policy, and Off-Grid Mode. ## Send photos to a vision model When the remote model row shows the eye icon, attach a photo as you would for a local vision model. Photo input works across Local Network, Private Link, Direct, Noema Bridge, and Cloud Relay. For a GGUF model, the Mac must have the required multimodal projector. A repository may describe the family as multimodal even when the downloaded weights do not include that projector. In that case Noema treats the installation as text-only and does not show the eye icon. > Current request limits are five images, up to 12 MB each and 40 MB total. ## Download a copy instead Long-press or open the model row’s context menu and choose **Download to this device**. Downloading creates a local installation; it is different from using the Mac’s existing copy through Constellation. ## Related documentation - [Noema Constellation](https://noemaai.com/docs/constellation) - [Constellation Connections & Privacy](https://noemaai.com/docs/constellation-connections) - [Running Models Locally](https://noemaai.com/docs/running-llms) - [Chat Tools & Tool Store](https://noemaai.com/docs/tool-usage) - [Datasets & RAG](https://noemaai.com/docs/dataset-integration) --- --- title: "Remote Endpoints" description: "Connect Noema to OpenAI-compatible, LM Studio, Ollama, OpenRouter, or other intentional remote inference endpoints." version: "Noema 3.5+" platforms: ["iPhone", "iPad", "Mac", "Vision Pro"] reviewed: "July 20, 2026" canonical: "https://noemaai.com/docs/remote-endpoints" --- # Remote Endpoints Provider disclosure, endpoint security, supported routes, credentials, and Off-grid behavior. ## Remote inference sends chat data to the provider _Data disclosure_ When you select a remote model, its provider may receive the current prompt, relevant conversation history, system and chat instructions, supported attachments, tool definitions and results, and retrieved excerpts needed for the request. The provider’s retention, training, security, and account policies apply. Use a remote endpoint only when that disclosure is intentional. ## Supported endpoint styles | Type | Behavior | | --- | --- | | OpenAI-compatible | Uses the configured base URL and compatible chat or response routes. | | LM Studio | Can use LM Studio’s native discovery and model-loading routes in addition to compatible chat APIs. | | Ollama | Discovers and addresses models through Ollama’s local server routes. | | OpenRouter | Uses a hosted catalog and provider routing with an API key. | | Custom | Uses the route, headers, and model identifiers you supply. | ## Secure the connection - Prefer HTTPS with a valid certificate for traffic that leaves your device. - On a LAN, bind the server deliberately, restrict the firewall, and do not expose an unauthenticated endpoint to the public internet. - Treat plain HTTP as readable by other systems on that network. - API secrets are stored in Keychain where the endpoint integration supports credential storage; do not put secrets into model names or shared screenshots. ## Set up and test an endpoint 1. Open Stored and choose Add remote endpoint. 2. Select a provider type and enter the base URL, model identifiers, and any required credential. 3. Test discovery and a small prompt before using it for sensitive work. 4. Review per-model context and sampling controls exposed by that provider. 5. Watch the connection state and response receipt for route or compatibility errors. ## Off-grid behavior Off-grid Mode blocks remote HTTP and HTTPS inference. Noema keeps the request local when a compatible local fallback exists; otherwise the turn explains that the selected endpoint is unavailable. ## Related documentation - [Privacy & Network Activity](https://noemaai.com/docs/privacy-and-network-activity) - [Noema Autopilot](https://noemaai.com/docs/autopilot) - [Noema Constellation](https://noemaai.com/docs/constellation) --- --- title: "Chat Interface" description: "Use Noema chat modes, document attachments, status controls, citations, receipts, diagnostics, context plans, and export packs." version: "Noema 3.6+" platforms: ["iPhone", "iPad", "Mac", "Vision Pro"] reviewed: "July 26, 2026" canonical: "https://noemaai.com/docs/chat-interface" --- # Chat Interface Current conversation controls, message actions, and route visibility, including Constellation sessions. ## Before you send - Choose a local model, remote model, Constellation model, or Autopilot. - Select a chat mode and answer style, add chat instructions, and set reasoning control when the model supports it. - On iPhone, iPad, or Mac, attach PDFs from the Files interface and wait for the inline indexing status to finish before asking about them. Supported EPUB and text attachments can also be indexed, and you can select an existing dataset for retrieval. - Arm only the chat tools you want available for the next turn. - Open the chat status drawer to review capabilities, reasoning, Autopilot, MCP, and the current context budget before sending. ## Read the chat status drawer - Compact capability controls keep the active chat features together. - Reasoning and Autopilot status show how the next turn is configured. - MCP visibility makes connected tool availability easier to inspect. - The detailed context-budget meter shows how much of the model’s window is planned or occupied. ## Keep longer conversations responsive As a conversation grows, Noema can compact older context so the active thread remains useful without making every turn carry the full history unchanged. This reduces the chance that a long chat stalls as it approaches the model’s context limit. > Compaction preserves room for the current exchange, but it is still a condensed representation of earlier context. Start a focused chat when exact wording from much earlier turns is essential. ## Message actions | Action | What it does | | --- | --- | | Branch | Starts an alternate continuation from the selected point. | | Regenerate | Runs the response again with the active route and settings. | | Try Model | Retries the turn with another compatible model. | | Bookmark | Saves the answer for later review. | | Pin to Private Scratchpad | Keeps selected material in the private scratchpad. | | Audit | Opens supporting runtime, citation, receipt, and diagnostic detail. | | Reroute | Retries an Autopilot answer through a different allowed route. | ## Find and review earlier work - Chat Recall searches past conversations when enabled. - Bookmarks collect saved responses without inventing folders or notebook collections. - Raw output exposes the model’s unprocessed response when you need to diagnose formatting. - Runtime information shows model, route, token, timing, and context details available for the turn. ## Evidence and receipts Web and dataset answers can include cleaner citations, evidence controls, retrieval receipts, and source detail. Numbers and complex formulas use sharper math rendering. Open the source when accuracy matters; a citation shows what material was provided, not that the model interpreted it correctly. ## Export and diagnose - Export packs collect the conversation and selected supporting material for handoff or review. - Diagnostics expose route failures, model events, retrieval behavior, and tool execution without relying on nonexistent message deletion or duplicate controls. ## Related documentation - [Chat Tools & Tool Store](https://noemaai.com/docs/tool-usage) - [Noema Autopilot](https://noemaai.com/docs/autopilot) - [Datasets & RAG](https://noemaai.com/docs/dataset-integration) --- --- title: "Datasets & RAG" description: "Import supported documents and media, download datasets, build local indexes, and inspect retrieval health in Noema." version: "Noema 3.5+" platforms: ["iPhone", "iPad", "Mac", "Vision Pro"] reviewed: "July 20, 2026" canonical: "https://noemaai.com/docs/dataset-integration" --- # Datasets & RAG Supported imports, embeddings, Hugging Face datasets, retrieval settings, and Dataset Health. ## What you can add manually | Material | Examples and limits | | --- | --- | | Documents | PDF, EPUB, Markdown, plain text, JSON, and JSONL, subject to extraction support. | | Media | M4A, MP3, WAV, AAC, AIFF, CAF, MOV, MP4, and M4V through a configured transcription engine. | | Tabular files | Direct manual CSV and TSV imports are currently rejected, even though downloaded dataset packages can contain them. | ## Other dataset sources - Download compatible Hugging Face datasets from Explore. - Install curated Knowledge Packs with source, license, and snapshot information. - Receive managed enterprise datasets when enrolled in a workspace. ## Embeddings are required for retrieval A chat model writes the answer; an embedding model builds and searches the semantic index. Install the embedding model before indexing documents, datasets, or Knowledge Packs. 1. Import or download the material. 2. Allow extraction or transcription to complete. 3. Configure chunking and retrieval settings when needed. 4. Wait for embeddings and the index to finish. 5. Select the dataset for a chat or arm Dataset Search. ## Dataset Health and maintenance - Use Dataset Health to inspect missing, stale, or incomplete index state. - Refresh when source material changed; rebuild when extraction, chunking, or embeddings need to be regenerated. - Tune retrieval count and relevance behavior to the model’s context budget. - Dataset Search can query indexed material beyond the dataset selected for the conversation. ## Citations require judgment Retrieved passages include source information and relevance signals. Noema can ask the model to cite them, but it cannot promise that every generated answer identifies the exact passage that determined the conclusion. Open the underlying evidence for high-stakes work. ## Related documentation - [Knowledge Packs](https://noemaai.com/docs/knowledge-packs) - [Voice & Media](https://noemaai.com/docs/voice-and-media) - [Privacy & Network Activity](https://noemaai.com/docs/privacy-and-network-activity) --- --- title: "Knowledge Packs" description: "Install curated offline Knowledge Packs and use their indexed source material in private, citation-aware chats." version: "Noema 3.1+" platforms: ["iPhone", "iPad", "Mac", "Vision Pro"] reviewed: "July 20, 2026" canonical: "https://noemaai.com/docs/knowledge-packs" --- # Knowledge Packs Curated, license-cleared offline collections for local retrieval. ## Available packs | Pack | Source | Important note | | --- | --- | --- | | Wilderness & Survival | U.S. Army survival and land-navigation field manuals | Verify actions involving environmental or medical risk. | | First Aid & Field Medicine | U.S. Army first-aid material | Not a substitute for professional care or emergency services. | | Emergency Preparedness | Ready.gov and FEMA | June 2026 snapshot; follow current local emergency instructions. | | Travel — World Factbook | CIA World Factbook | Early-2026 snapshot; changing facts may be outdated. | ## Install a pack 1. Open Explore → Datasets and find Knowledge Packs. 2. Review the pack’s description, source, license, snapshot date, and indexing estimate. 3. Download the pack and install the embedding model if prompted. 4. Wait until indexing completes before selecting it in chat. ## Use a pack in chat Select the pack as the active dataset or arm Dataset Search with a tool-capable model. Retrieved passages include source information and relevance scores, but you should still open the underlying text for important decisions. ## Offline and Autopilot behavior Downloading a pack requires network access. Indexing, semantic search, and local-model generation happen on-device afterward. Knowledge Pack excerpts stay local in Autopilot by default. They are sent to a remote stronger model only when you explicitly allow escalation for knowledge-base chats. ## Refresh or delete Use Dataset Health to rebuild stale or incomplete extraction, chunks, and embeddings. Deleting a pack removes its downloaded source files and local retrieval index; it can be installed again later. ## Related documentation - [Datasets & RAG](https://noemaai.com/docs/dataset-integration) - [Quick Start](https://noemaai.com/docs/quick-start) - [Privacy & Network Activity](https://noemaai.com/docs/privacy-and-network-activity) --- --- title: "Voice & Media" description: "Use hands-free Voice Mode, configure local transcription and speech, and turn audio or video files into searchable material." version: "Noema 3.1+" platforms: ["iPhone", "iPad", "Mac", "Vision Pro"] reviewed: "July 20, 2026" canonical: "https://noemaai.com/docs/voice-and-media" --- # Voice & Media Speech recognition, spoken responses, Voice Mode, and media transcription. ## Choose a transcription engine | Engine | Data path | | --- | --- | | Apple Speech | May use Apple services unless on-device recognition is available and required. | | WhisperKit | Runs a downloaded Whisper model locally with Apple-optimized inference. | | whisper.cpp | Runs compatible Whisper models locally through the native runtime. | | Audio-language model | Sends the required media to a configured remote audio endpoint. | > Open Settings → Speech & ASR. Off-grid Mode requires an on-device transcription path. ## Voice Mode The transcript and generated response remain in the current conversation. 1. Noema listens for speech and the selected ASR engine produces a transcript. 2. The active chat model generates a response. 3. The selected voice engine speaks the answer. 4. The session returns to listening until you end Voice Mode. ## Voice output - Neural Voice uses Noema’s optional local neural voice model on supported Apple-silicon devices with sufficient memory. - System Voice uses operating-system speech synthesis and exposes speaking-rate control. - If Neural Voice is unavailable or its model is missing, Noema falls back to System Voice. ## Audio and video imports Supported media includes M4A, MP3, WAV, AAC, AIFF, CAF, MOV, MP4, and M4V. Noema transcribes spoken content so the result can be reviewed in chat or indexed in a local dataset. ## Troubleshooting - Check microphone and speech-recognition permissions. - Confirm local Whisper or Neural Voice model downloads are complete. - Keep enough free storage and memory for long media. - Review names, numbers, and specialized terminology in every transcript. - Off-grid Mode blocks remote speech and audio endpoints. ## Related documentation - [Datasets & RAG](https://noemaai.com/docs/dataset-integration) - [Privacy & Network Activity](https://noemaai.com/docs/privacy-and-network-activity) - [Support](https://noemaai.com/docs/support) --- --- title: "Python Code Execution" description: "Run CPython locally in Noema with bundled packages for data analysis, machine learning, symbolic mathematics, graph analysis, and scientific visualization." version: "Noema 3.7+" platforms: ["iPhone", "iPad", "Mac", "Vision Pro"] reviewed: "July 30, 2026" canonical: "https://noemaai.com/docs/python-code-execution" --- # Python Code Execution Run Python locally inside Noema for calculations, data analysis, machine learning, symbolic mathematics, graph analysis, and scientific visualization. ## Overview The Python tool executes code using the CPython 3.14.2 runtime bundled with Noema. It is designed for offline data analysis and scientific computing without sending your code or data to a remote Python service. The expanded package set documented below is available in Noema 3.7 and later. - Perform numerical and statistical calculations. - Analyze tables and CSV files with pandas. - Build machine-learning models with scikit-learn. - Create charts with Matplotlib and seaborn. - Use SciPy for optimization, integration, statistics, signal processing, and linear algebra. - Perform symbolic mathematics with SymPy. - Analyze graphs and networks with NetworkX. - Read and write YAML data. - Format results as readable tables. > The runtime is restricted for safety. It cannot access the network, launch other processes, or install additional packages. ## Enabling Python Python must be enabled globally and turned on for the current conversation before a model can use it. 1. Open **Settings**. 2. Select **Tools**. 3. Enable **Python**. 4. Open a conversation. 5. Use the tool toggles in the context bar to turn Python on for that conversation. The selected model must support tool calling. Noema will show Python as unavailable if the embedded runtime fails its health check. Open the Python page in Settings under Tools to see the installed Python version, package-set identifier, package health, and the complete package list with versions and dependencies. ## Bundled packages The following packages are included with Noema: | Package | Version | Import name | Primary uses | | --- | --- | --- | --- | | NumPy | 2.5.1 | `numpy` | Arrays, numerical computing, random numbers, and linear algebra | | pandas | 3.0.5 | `pandas` | DataFrames, CSV processing, grouping, dates, and tabular analysis | | Matplotlib | 3.11.1 | `matplotlib` | Scientific charts, subplots, annotations, and custom figures | | SciPy | 1.18.0 | `scipy` | Statistics, optimization, integration, FFTs, and scientific algorithms | | scikit-learn | 1.9.0 | `sklearn` | Preprocessing, regression, classification, clustering, and model evaluation | | Pillow | 12.3.0 | `PIL` | Image creation and processing | | seaborn | 0.13.2 | `seaborn` | Statistical visualization built on Matplotlib | | SymPy | 1.14.0 | `sympy` | Symbolic algebra, calculus, equations, and exact mathematics | | NetworkX | 3.6.1 | `networkx` | Graph algorithms and network analysis | | PyYAML | 6.0.3 | `yaml` | Reading and writing YAML | | tabulate | 0.10.0 | `tabulate` | Plain-text and Markdown table formatting | ## Bundled package dependencies Noema also bundles the dependencies required by these packages: | Dependency | Version | | --- | --- | | contourpy | 1.3.3 | | cycler | 0.12.1 | | fontTools | 4.63.0 | | joblib | 1.5.3 | | kiwisolver | 1.5.0 | | mpmath | 1.3.0 | | narwhals | 2.24.0 | | packaging | 26.2 | | pyparsing | 3.3.2 | | python-dateutil | 2.9.0.post0 | | six | 1.17.0 | | threadpoolctl | 3.6.0 | | wcwidth | 0.8.2 | Package versions may change with Noema updates. The Python help sheet inside the app shows the exact package set installed by your current build. Commands such as `pip install` are unsupported. Native Python extensions are compiled, signed, and included with the application. > Packages are bundled with Noema. Additional packages cannot be installed at runtime. ## Creating Matplotlib charts Noema uses Matplotlib’s non-interactive Agg backend. Interactive desktop windows and GUI backends are unavailable, but completed figures can be displayed directly inside the conversation. You can display a figure by calling `plt.show()` or by leaving it open when execution finishes. Noema automatically captures open figures even if later code reports an error. - Captured charts appear directly below the completed Python tool call. - Charts can be opened in a zoomable viewer. - Charts can be shared or saved from the viewer. - Charts preserve their aspect ratio across iPhone, iPad, Mac, and Vision Pro. - Inline previews use PNG and are retained with the conversation. Noema captures up to four figures from one execution. Each inline preview is limited to 1600 × 1200 pixels and 700 KB, with a combined preview limit of 2.8 MB. Explicitly saved PNG and JPEG files can also be returned as regular artifacts. SVG and PDF figures are not currently displayed inline. For example, ask: “Use NumPy and Matplotlib to create a sine and cosine chart with labeled axes, a legend, and a grid. Display it inline and save it as `trigonometry.png`.” ## Python or Quick Charts? | Use Quick Charts when… | Use Python and Matplotlib when… | | --- | --- | | The values are already calculated | The data must be transformed or analyzed first | | You need a simple bar, line, scatter, or pie chart | You need histograms, heatmaps, regressions, or scientific axes | | You want a fast native chart | You need annotations, custom styling, or multiple subplots | | Python is disabled or unavailable | You need NumPy, pandas, SciPy, or scikit-learn integration | Noema should use only one visualization path for a requested chart. If you explicitly request Quick Charts or Matplotlib, the model should respect that choice. ## Execution restrictions Python runs locally, but it is restricted execution rather than a process-isolated security sandbox. - Has no network access. - Cannot create sockets or make HTTP requests. - Cannot launch subprocesses or shell commands. - Cannot fork, spawn, or use multiprocessing process backends. - Cannot load arbitrary dynamic libraries through `ctypes`. - Can load only the native extensions bundled and signed with Noema. - Can read the bundled Python runtime, installed packages, the current execution directory, and approved font directories. - Can write only to the execution directory and Noema’s bounded Matplotlib cache. - Limits numerical thread pools to one thread by default. - Has a 30-second execution timeout. Code that attempts a prohibited operation will receive an error without weakening the restrictions for the rest of the runtime. ## Results and files A Python execution can return: - Standard output produced by `print()`. - Standard error. - Exit status and execution time. - Timeout or runtime errors. - Matplotlib figures. - Files created in the execution directory. Use `print()` when you want textual results to appear clearly in the tool response. Large binary data is not sent back to the language model, although supported artifacts can remain available to you in the app. ## Availability with models and datasets Python requires a model capable of calling tools. On-device Apple Foundation Models do not currently execute the Python tool. Private Cloud Compute models can use it when tool calling is available. For most inference backends, Python is unavailable while a dataset is selected or being indexed. Constellation can coordinate dataset retrieval and Python execution on the originating device where both capabilities are available. ## Example prompts - Use NumPy to generate 1,000 normally distributed values and report their mean, standard deviation, and percentiles. - Create a pandas DataFrame containing monthly revenue and expenses, calculate profit and growth, and print the result as a Markdown table. - Use SciPy to minimize the function `(x - 3)² + 2` and explain the result. - Train a small scikit-learn linear regression model on sample data, show its coefficients, and plot the fitted line. - Use seaborn and Matplotlib to create a correlation heatmap for a sample dataset and display it inline. - Use SymPy to solve `x³ - 6x² + 11x - 6 = 0` exactly and verify each solution. - Use NetworkX to create a graph, calculate PageRank and shortest paths, and summarize the most important nodes. - Create a YAML document containing a project name, milestones, and owners, then parse it back and print the result. ## Troubleshooting - If Python is unavailable, confirm that it is enabled in Settings, turned on using the tool toggles in the context bar, and supported by the selected model. Also check whether a dataset is currently selected. - If a bundled package fails to import, open the Python help sheet in Settings and check the runtime and package health diagnostics. - If a chart does not appear, call `plt.show()` or leave the figure open. Interactive Matplotlib windows are not supported. - If execution times out, reduce the dataset size, number of iterations, image resolution, or model complexity. Multiprocessing cannot be used to bypass the execution limit. - If code attempts to download data or install a package, provide the required data as a local attachment or rewrite the task using the packages already bundled with Noema. ## Related documentation - [Chat Tools & Tool Store](https://noemaai.com/docs/tool-usage) - [Constellation Models, Tools & Photos](https://noemaai.com/docs/constellation-models-and-tools) - [Privacy & Network Activity](https://noemaai.com/docs/privacy-and-network-activity) - [Support](https://noemaai.com/docs/support) --- --- title: "Private Translation in Noema" description: "Translate text privately on your Apple device with a RAM-aware Hy-MT2 model, automatic language detection, and formatting preservation." version: "Noema 3.7+" platforms: ["iPhone", "iPad", "Mac", "Vision Pro"] reviewed: "August 17, 2026" canonical: "https://noemaai.com/docs/translation" --- # Private Translation in Noema Download a translation model once, then translate among 38 languages without sending the text to a remote AI service. Noema can translate text directly on your device with a local Tencent Hy-MT2 model. After the model is downloaded, the translation itself does not require an internet connection and does not send the source text to a remote AI provider. The Translation workspace supports 38 languages, automatic source-language detection, one-tap language swapping, optional formatting preservation, and paragraph-by-paragraph processing for longer text. Noema also checks the device's current app memory budget before showing or downloading a model, so it does not recommend a translation model that it considers unsafe to load. > Translation is generated by an AI model and can contain errors. Review names, numbers, technical terms, legal language, medical information, and other high-stakes content before relying on the result. ## Quick start 1. Open **Tools** in Noema. 2. Under **Language**, open **Translation**. 3. If prompted, download a compatible Hy-MT2 model. Noema only shows models that fit the device's current memory budget. 4. Choose a **Source language**, or leave it on **Detect language**. 5. Choose the **Target language**. 6. Enter text in **Original**, or use the clipboard button to paste it. 7. Leave **Preserve formatting** on when line breaks, markup, code, or placeholders matter. 8. Select **Translate**. With a hardware keyboard, you can also press **Command–Return**. 9. When the translation is complete, select **Copy** to place it on the clipboard. You can cancel a translation while it is running. For multi-paragraph text, Noema shows paragraph progress in Simple and Advanced modes. ## First-time setup Translation uses a dedicated local model. The first time you open the tool, Noema evaluates the device's current app memory budget and offers only compatible options. In **Beginner** mode, setup is simplified: - Noema chooses the smallest Hy-MT2 model that passes the memory check. - The setup page shows the one-time download size before downloading. - The download can be paused, resumed, cancelled, or retried. - Noema checks available storage and waits for a usable connection when necessary. In **Simple** and **Advanced** modes, you can choose among every model that currently fits. Noema marks the largest fitting option as **Recommended** for quality, while keeping smaller models available for faster translation and smaller downloads. Noema recalculates the list from the device's reported app memory budget and checks the fit again immediately before starting a download. ## Translation models Noema supports three official Hy-MT2 GGUF options. All three use the `Q4_K_M` quantization supported by Noema's local llama.cpp runtime. | Model | Approximate download | Noema's guidance | Important detail | | --- | --- | --- | --- | | [Hy-MT2 1.8B](https://huggingface.co/tencent/Hy-MT2-1.8B-GGUF) | 1.13 GB | Fastest; suited to everyday translation | The automatic choice in Beginner mode when it fits | | [Hy-MT2 7B](https://huggingface.co/tencent/Hy-MT2-7B-GGUF) | 4.62 GB | Higher quality; better nuance and instruction following | Requires substantially more storage and memory than 1.8B | | [Hy-MT2 30B-A3B](https://huggingface.co/tencent/Hy-MT2-30B-A3B-GGUF) | 18.24 GB | Maximum quality for complex translation | Currently runs CPU-only in Noema and can be considerably slower | Download size is not the same as runtime memory use. Noema's fit check includes model weights, an 8,192-token context, the KV cache, runtime buffers, and conservative working headroom. A model can therefore be hidden even when the device has enough free storage for its file. There is no universal “best” model for every device. Start with 1.8B when responsiveness matters, use 7B when the device can fit it and nuance matters more, and reserve 30B-A3B for memory-rich devices and work where its much slower CPU-only execution is acceptable. ## Supported languages Any supported language can be selected as the source or target. The source can also be detected automatically. - Arabic, Bengali, Burmese, Cantonese - Chinese, Traditional Chinese, Czech, Dutch - English, Filipino, French, German - Gujarati, Hebrew, Hindi, Indonesian - Italian, Japanese, Kazakh, Khmer - Korean, Malay, Marathi, Mongolian - Persian, Polish, Portuguese, Russian - Spanish, Tamil, Telugu, Thai - Tibetan, Turkish, Ukrainian, Urdu - Uyghur, Vietnamese Language names follow the language selected for Noema's interface when a localized system name is available. Automatic detection works best when the source is primarily one language. For short, ambiguous, or mixed-language text, choosing the source language explicitly can improve consistency. Noema disables translation when the explicitly selected source and target are the same. ## Preserving formatting Turn on **Preserve formatting** when the structure of the source matters. The model is instructed to preserve: - Line breaks and paragraph boundaries - Delimiters and placeholders - Markup - Code - Other non-user-facing structure Noema separates the source at blank-line paragraph boundaries, translates each paragraph independently, and then restores the original blank-line separators exactly. A translation is shown only after every paragraph has returned a usable result, which prevents a partially translated document from being presented as complete. Formatting within a translated paragraph is still model-generated. Always verify code, markup, variables, URLs, product names, and placeholders before using the result in production. Turn **Preserve formatting** off when you want the model to use more natural formatting in the target language. ## Translating longer text Noema processes multi-paragraph text one paragraph at a time. This has several benefits: - Long documents do not have to fit into the model as one request. - Existing blank-line structure can be retained. - Progress can be reported paragraph by paragraph. - A malformed paragraph can be retried without accepting an incomplete final document. Each paragraph still has to fit inside the translation session's bounded context. If Noema reports that the text is too long for the current context, divide an unusually long unbroken paragraph into smaller paragraphs with blank lines, then translate again. The Translation workspace currently accepts text entered or pasted into the editor. It does not directly import a PDF, Word document, image, or webpage. Extract or copy the text first, then review the translated formatting before placing it back into the original document. ## Automatic detection and language swapping Choose **Detect language** when you do not know the source language. Each translated paragraph reports a detected source language, and Noema uses the most frequently detected language for the completed result. After a translation finishes, select the swap button to: - Make the previous target language the new source language - Use the detected or selected source language as the new target - Move the translated output into the **Original** editor This is useful for translating a result back in the opposite direction. A round-trip translation is a convenience check, not proof that either translation is exact. ## Privacy and offline use The Translation workspace only accepts a recognized local Hy-MT2 GGUF model. It does not route translation text through a configured remote endpoint, a cloud model, or the chat model's provider. - **During model setup:** Noema connects to Hugging Face to download the selected Hy-MT2 model file. - **During translation:** The text is processed by Noema's embedded local llama.cpp runtime on the device. - **After setup:** Translation can run without an internet connection, including in Off-Grid mode. - **Chat history:** Translation requests and results are not inserted into a chat conversation. The workspace does not provide a saved translation history. Copy any result you want to keep. ## How translation affects the chat model Noema treats Translation as an isolated workspace. It may temporarily load the selected Hy-MT2 model into the local inference runtime, but it does not replace your normal chat-model preference. - In Simple and Advanced modes, the translation model remains available while you stay in the workspace. Your previous chat model is restored when you leave. - In Beginner mode, Noema returns to the managed primary model after a translation finishes or is cancelled. - Entering Beginner mode while a translation is active cancels the translation and applies Beginner's managed runtime policy. Loading and restoring large local models can take time. This is expected and is separate from the translation itself. ## Tips for better translations - Select the source language instead of automatic detection for very short text, names, or multilingual content. - Keep **Preserve formatting** on for templates, localization strings, code-adjacent text, or content with placeholders. - Turn it off for ordinary prose when natural target-language formatting matters more than source layout. - Split extremely long paragraphs at natural boundaries. - Use the 7B model when nuance matters and the device offers it safely. - Use the 1.8B model when speed and lower memory use matter most. - Treat the 30B-A3B model as a quality-focused, CPU-only option rather than a fast default. - Check terminology, proper nouns, measurements, dates, numbers, and negation manually. - Keep an untouched copy of the source when translating structured content. ## Troubleshooting | What you see | What it means | What to do | | --- | --- | --- | | **No Hy-MT2 model fits safely** | Even the smallest supported model is above Noema's reported app memory budget. | Select **Check Again** to recalculate the fit. If no model appears, private translation is not available under that device's reported budget. | | **Private translation isn't available on this device** | Beginner mode could not find a compatible translation model that safely fits. | Select **Check Again**. Use another supported device if the result does not change. | | **Free some storage, then try again** | The one-time model download does not have enough working storage. | Free local storage and retry. The required free space can be larger than the final model file while a download is in progress. | | **Waiting for a connection** | The model is not installed and the download cannot currently proceed. | Reconnect, then select **Try Again**. On a constrained or metered connection, confirm the download when Noema asks. | | The **Translate** button is unavailable | The input is empty, no compatible model is installed, the target is missing, the explicit source equals the target, or a translation is already running. | Check the text, language pair, model setup, and current progress. | | **This text is too long for the translation model's current context** | At least one paragraph cannot fit in the bounded translation context. | Break the long paragraph into smaller paragraphs with blank lines and retry. | | **Failed to load the translation model** | The local runtime could not load the selected model and settings. | Free memory, retry, or choose a smaller model from **Models**. | | **The model returned an unreadable translation** | The model did not return a complete structured paragraph, even after Noema retried it once. | Try again. If it repeats, shorten the paragraph or choose another fitting model. | | The result changed markup or placeholders | Formatting preservation is model-assisted and is not a formal parser. | Turn on **Preserve formatting**, translate smaller sections, and compare the result with the source before use. | | Translation feels slow | Model loading and local generation depend on the model and device. The 30B-A3B option is CPU-only. | Keep the workspace open between translations or choose a smaller model. | ## Current limitations - Translation is text-only inside this workspace. - The tool does not directly import document files, images, audio, or webpages. - Only the listed 38 languages are available in the language selectors. - There is no built-in glossary, translation memory, or terminology lock. - There is no persistent translation history. - Formatting preservation cannot guarantee byte-for-byte preservation inside model-generated paragraphs. - Automatic detection can be ambiguous for short or mixed-language input. - The 8,192-token context applies per paragraph, not to an unlimited single block of text. - Translation quality varies by language pair, model size, subject, and writing style. ## Frequently asked questions ### Does translation require an internet connection? Only for the model download. Once a compatible Hy-MT2 model is installed, the translation itself runs locally and can work offline. ### Is my text sent to Tencent or Hugging Face? Noema contacts Hugging Face to download the model file. The Translation workspace then runs that model locally; it does not send the text to Tencent, Hugging Face, or a configured remote inference provider. ### Why do I see fewer model choices than another device? Noema filters the catalog using the current device's app memory budget and a conservative runtime estimate. A device with a larger safe budget can show larger models. ### Why is the largest model not always available? The 30B-A3B download is approximately 18.24 GB, requires much more runtime memory than its file size alone suggests, and currently runs CPU-only in Noema. It only appears when the device passes the memory-fit check. ### Which model should I choose? Choose 1.8B for speed and broad device compatibility, 7B for stronger nuance when it fits, and 30B-A3B only when maximum quality matters more than speed. Noema's **Recommended** badge identifies the largest option that currently passes its safety check. ### Can Translation use my OpenAI, OpenRouter, Ollama, LM Studio, or Constellation model? No. This workspace deliberately uses one of Noema's recognized local Hy-MT2 GGUF models. ### Does Translation change my default chat model? No. The Hy-MT2 runtime is temporary. Noema restores the applicable chat runtime when the Translation workspace ends, with additional managed restoration in Beginner mode. ### Are translations saved in my conversations? No. They are not inserted into chat history. Use **Copy** if you want to keep or move a result. ### Can I translate a whole document? You can paste multi-paragraph text, and Noema will process it paragraph by paragraph. Direct file import and layout-aware document export are not currently part of the Translation workspace. ### Can I trust the output without review? No machine translation should be assumed perfect. Review important output, especially content involving safety, health, law, money, identity, deadlines, or contractual obligations. ## How it works Noema's translation pipeline is designed to favor complete, inspectable results: 1. The app verifies that the selected model is a supported Hy-MT2 GGUF. 2. It divides the source into paragraphs at blank-line boundaries and assigns each paragraph a stable internal identifier. 3. It sends each paragraph to the local model with the chosen language direction and formatting policy. 4. It requires a structured result containing the matching paragraph identifier, the translation, and the detected source language. 5. It rejects empty, mismatched, malformed, or token-limited output and retries the paragraph once. 6. It assembles the completed paragraphs with the source's original blank-line separators. 7. It publishes the final translation only if every paragraph is present and non-empty. Source content is explicitly marked as content rather than instructions. The translation calls are also treated as auxiliary local work, so they do not create durable chat-generation checkpoints. ## Related documentation - [Downloading Models](https://noemaai.com/docs/downloading-models) - [Running Models Locally](https://noemaai.com/docs/running-llms) - [Privacy & Network Activity](https://noemaai.com/docs/privacy-and-network-activity) --- --- title: "Chat Tools & Tool Store" description: "Enable and arm Noema chat tools, review permissions, and distinguish callable tools from the separate Tools tab." version: "Noema 3.6+" platforms: ["iPhone", "iPad", "Mac", "Vision Pro"] reviewed: "July 26, 2026" canonical: "https://noemaai.com/docs/tool-usage" --- # Chat Tools & Tool Store The full callable tool catalog, permissions, confirmations, and per-chat controls. ## Enable globally, arm per chat 1. Open Settings → Tools and enable the capabilities you are willing to make available. 2. Grant the operating-system permissions required by Calendar, speech, files, or other integrations. 3. In Chat, open the tools menu and arm only the tools the current conversation should use. 4. Use a tool-capable model. A tool being installed does not force the model to call it. ## Callable chat tools | Tool | Purpose | | --- | --- | | Web | Research, open, and find across supported web evidence. | | Python | Local sandboxed calculations, parsing, and structured computation; compatible MLX models are not categorically excluded. | | Memory | Store or retrieve user-approved durable notes. | | Calculator | Deterministic arithmetic. | | Unit Converter | Deterministic measurement conversion. | | Dataset Search | Search indexed datasets and Knowledge Packs. | | PDF Reader | Read supported PDF text; it becomes available automatically when applicable. | | Calendar | Read calendar data and create events after confirmation and permission checks. | | Charts | Turn structured values into a visual chart. | | Phone a Friend | Let a local model request an Autopilot handoff when configured. | ## Permissions and confirmations Calendar writes require confirmation. File, speech, and calendar access remain subject to platform permissions. Remote tools or models may receive the arguments and results required for the turn; review Privacy & Network Activity before combining sensitive material with a remote route. ## Python is local and constrained - Runs in a temporary local workspace with execution limits. - Has no general network access. - Returns structured output so the model can inspect stdout, stderr, exit state, timing, and errors. - Availability depends on the current model’s tool-calling support and device runtime, not a blanket MLX prohibition. ## Chat tools are not the Tools tab Callable tools are functions the active chat model can invoke. The separate Tools tab contains user-facing workflows such as study and utility surfaces; opening that tab does not arm tools for a conversation. ## Tools with Constellation With a Constellation model, the model may run on your Mac while tool permission and execution remain on the device that started the chat. This lets local retrieval, PDF navigation, Memory, Python, Calendar, charts, Web Search, and selected MCP tools participate without moving the entire local library to the Mac. ## Related documentation - [Python Code Execution](https://noemaai.com/docs/python-code-execution) - [Constellation Models, Tools & Photos](https://noemaai.com/docs/constellation-models-and-tools) - [Web Search](https://noemaai.com/docs/web-search) - [Noema Autopilot](https://noemaai.com/docs/autopilot) - [Chat Interface](https://noemaai.com/docs/chat-interface) --- --- title: "Noema Autopilot" description: "Configure Noema Autopilot to route each request between an everyday local model and a stronger local or remote model." version: "Noema 3.5+" platforms: ["iPhone", "iPad", "Mac", "Vision Pro"] reviewed: "July 20, 2026" canonical: "https://noemaai.com/docs/autopilot" --- # Noema Autopilot Per-message routing, Phone a Friend, escalation policy, privacy, fallbacks, and receipts. ## Two routing systems Smart Router can use Apple’s on-device Foundation Model, keeping the routing decision local, or a small configured remote model that receives the message and limited routing context. | System | How it decides | | --- | --- | | Smart Router | Evaluates each message before generation and selects the everyday or stronger model. | | Phone a Friend | Lets the local tool-capable model begin the turn and request a handoff only when needed. | ## Set up Autopilot 1. Open Settings → Autopilot and choose Smart Router or Phone a Friend. 2. Select the everyday local chat model. 3. Choose a configured stronger remote model or, on supported Macs, a second installed local model that fits alongside it. 4. For Smart Router, select an on-device or remote router. 5. Review the privacy disclosure, test the configuration, and turn on Autopilot. ## Escalation levels | Level | Behavior | | --- | --- | | Conserve | Keeps nearly everything on-device. | | Balanced | Escalates when the stronger model is likely to make a meaningful difference. | | Frontier | Escalates more readily when answer quality may benefit. | ## Privacy controls Pause cloud escalation keeps answers on-device. With a remote Smart Router, that router may still evaluate the message even though the answer is not handed to the stronger cloud model. Knowledge-base excerpts stay local by default. They are included in a remote escalation only when you enable Allow escalation for knowledge-base chats. ## Failure behavior and receipts If routing is unavailable, times out, or produces an unreadable decision, Noema falls back to a local heuristic. If escalation is blocked by Off-grid Mode, unavailable, or incompatible with an attachment, the turn stays local. - The response can show where the answer ran and why. - Receipts identify local fallbacks, the stronger model, and user overrides. - Totals can summarize local, cloud, and override behavior together with estimated energy or remote-cost savings. - Use message actions to retry or reroute when you want a different model. > Autopilot improves model selection; it does not guarantee factual accuracy. Review citations and source material for important decisions. ## Related documentation - [Privacy & Network Activity](https://noemaai.com/docs/privacy-and-network-activity) - [Remote Endpoints](https://noemaai.com/docs/remote-endpoints) - [Chat Interface](https://noemaai.com/docs/chat-interface) --- --- title: "Web Search" description: "Configure Noema Web Search, understand research, open, and find, and review what the selected service receives." version: "Noema 3.5+" platforms: ["iPhone", "iPad", "Mac", "Vision Pro"] reviewed: "July 20, 2026" canonical: "https://noemaai.com/docs/web-search" --- # Web Search Search controls, source reading, custom SearXNG behavior, and privacy boundaries. ## Enable and arm Web Search 1. Open Settings → Tools and enable Web Search globally. 2. In Chat, open the tools menu and arm Web Search for the current conversation. 3. Ask for current information or explicitly request research when the model should search. ## Three web actions | Action | Purpose | | --- | --- | | research | Searches and assembles candidate evidence for a question. | | open | Reads a selected source when supported. | | find | Locates relevant text within an opened source. | > Text-layer PDFs can be read. The source reader does not execute page JavaScript or perform OCR on image-only material. ## What may be sent The selected search service receives the generated query, locale, and search controls. Rich retrieval can also send candidate URLs, snippets or source references, and open or find parameters to Noema’s source reader. Returned evidence is then provided to the active model. A custom SearXNG instance supplies snippet-only search results; it does not provide Noema’s richer server-side source reading. ## Availability and quotas Noema does not impose an in-app web-search quota or subscription. Availability can still be affected by the selected service, upstream providers, rate limits, maintenance, or network conditions. ## Verify the evidence - Open cited sources and check the publication date. - Use find to inspect the passage behind a claim. - Treat snippets as leads, not complete evidence. - Off-grid Mode blocks Web Search. ## Related documentation - [Chat Tools & Tool Store](https://noemaai.com/docs/tool-usage) - [Privacy & Network Activity](https://noemaai.com/docs/privacy-and-network-activity) - [Chat Interface](https://noemaai.com/docs/chat-interface) --- --- title: "Constellation Connections & Privacy" description: "Understand Local Network, Private Link, Direct, Noema Bridge, and Cloud Relay, including how Noema chooses a route and what each route means for privacy." version: "Noema 3.6+" platforms: ["iPhone", "iPad", "Mac", "Vision Pro"] reviewed: "July 26, 2026" canonical: "https://noemaai.com/docs/constellation-connections" --- # Constellation Connections & Privacy Noema tests the available paths when you press Use and selects the best working route before the first message. ## Route order | Route | When Noema uses it | What it means | | --- | --- | --- | | Local Network | Both devices can reach each other on the same LAN. | The request goes directly to the Mac on the local network and usually has the lowest latency. | | Private Link | You configured a reachable Tailscale, ZeroTier, or WireGuard address. | Traffic travels over your private mesh without router port forwarding. | | Direct | The devices can establish an encrypted peer-to-peer path across networks. | CloudKit introduces the devices, but no relay carries the conversation. | | Noema Bridge | A restrictive NAT or firewall prevents Direct. | End-to-end encrypted bytes pass through Noema’s relay; the relay cannot read the request or reply. | | Cloud Relay | Faster routes are unavailable. | Requests and replies use your private CloudKit database. It is the broadest-compatibility and slowest fallback. | Direct is not guaranteed. Carrier-grade NAT, hotel Wi-Fi, enterprise firewalls, and some mobile networks intentionally prevent peer-to-peer connectivity. Seeing Noema Bridge on 5G can therefore be normal rather than a configuration problem. ## See the current route - **Stored → Constellation**, under Remote Session. - The Chat connection button. - The menu containing the active model’s eject action. Constellation settings also explain every route. Private Link appears in the route list after one has been configured. ## Set up Private Link 1. Connect the Mac and requesting device to the same Tailscale, ZeroTier, or WireGuard network. 2. Open **Stored → Constellation** on the requesting device. 3. Open the Mac under the lower **Private Link** section. 4. Enter the private address assigned to the Mac. 5. Save, then press **Use** on a Mac model to test the route. No router port forwarding is required. Noema checks the private address directly and avoids routing that check through a system proxy. ## What Noema Bridge can see Noema Bridge receives an opaque routing token and encrypted traffic. The two devices authenticate the session and encrypt the request and response end-to-end. The bridge forwards sealed bytes and does not run the model or store its weights. Cloud Relay is a separate fallback. It stores the exchange in the user’s private CloudKit database rather than sending it through the Bridge socket. ## Remote Access and Off-Grid Mode Remote Access is off by default. Enabling it is a deliberate action on both the Mac and requesting device. Off-Grid Mode takes priority over Constellation. It pauses iCloud sync, discovery, Remote Access, and network-backed tools until the mode is turned off. ## Related documentation - [Noema Constellation](https://noemaai.com/docs/constellation) - [Constellation Models, Tools & Photos](https://noemaai.com/docs/constellation-models-and-tools) - [Privacy & Network Activity](https://noemaai.com/docs/privacy-and-network-activity) - [Privacy FAQ](https://noemaai.com/docs/privacy-faq) --- --- title: "Constellation Model Ejection & Help" description: "Fix device discovery, route, model-loading, vision, tool, sync, and idle-ejection issues in Constellation." version: "Noema 3.6+" platforms: ["iPhone", "iPad", "Mac", "Vision Pro"] reviewed: "July 26, 2026" canonical: "https://noemaai.com/docs/constellation-troubleshooting" --- # Constellation Model Ejection & Help Check both devices, confirm the Mac is awake, inspect the selected route, and use the ejection controls to manage Mac memory. ## Release the Mac model On the requesting device, open the active connection control in Chat and choose the eject action. Under **Settings → Sync & Devices → Model Ejection**, choose whether this also unloads the model from the connected device and whether that happens immediately or after a delay. On the Mac, open **Settings → Sync → Model Ejection** to control whether a requester may unload the Mac model, whether idle Constellation models unload automatically, and how long an idle model remains in memory. > Both devices must allow remote ejection. An active generation is never interrupted; Noema waits until it finishes before unloading the model. ## The Mac does not appear - Confirm both devices use the same Apple Account and can access iCloud. - Turn on **Reach this Mac from anywhere** under Mac **Settings → Sync**. - Turn on **Use your Mac’s models** under Constellation on the requesting device. - Keep Noema open and the Mac awake. - Turn off Off-Grid Mode. - Pull down in the iOS chat drawer, then reopen Stored. ## The Mac appears asleep Noema cannot wake a sleeping Mac, and closing a MacBook lid still puts it to sleep. Under Mac **Settings → Sync → Keep this Mac awake**, choose **While plugged in** or **Always** if that matches your power needs. ## Use is taking time Pressing **Use** starts the real model load immediately. Chat can appear before a large model is fully resident, but the first request must still wait for loading to finish. A model already marked loaded should start faster. ## The route is Bridge instead of Direct This usually means the current NAT or firewall rejected a peer-to-peer path. Bridge is the intended encrypted fallback. Configure Private Link if you operate a private mesh and want a predictable path across networks. ## The eye icon is missing - Confirm the selected model is actually a vision model. - For GGUF, install its matching `mmproj` or supported embedded projector beside the weights. - Reload the model after changing the projector. Noema hides vision capability when the installed files cannot process images, even if repository metadata describes the model family as multimodal. ## Tools do not appear or run - Open the Chat capability controls and allow the required tool. - Confirm the tool is enabled globally. - Check that required local data, runtime, or OS permission is available. - For MCP, select a connected server on a platform that supports it. - Confirm workspace policy and Off-Grid Mode permit the tool. The remote model proposes tool calls, but the requesting device owns permissions and execution. ## A chat has not updated yet Pull down in the iOS chat drawer to force a complete sync. If it still does not arrive, check the Sync status on both devices and confirm neither device reports a local history-write failure or a newer-version compatibility hold. ## A project is present but its PDF is unavailable Project definitions and source references sync; source files and dataset indexes do not. Import or download the corresponding source on the current device before using local retrieval there. ## Related documentation - [Noema Constellation](https://noemaai.com/docs/constellation) - [Constellation Connections & Privacy](https://noemaai.com/docs/constellation-connections) - [Constellation Models, Tools & Photos](https://noemaai.com/docs/constellation-models-and-tools) - [Support](https://noemaai.com/docs/support) --- --- title: "Privacy & Network Activity" description: "See which Noema features stay on-device, which use the local network, and what optional online services receive." version: "Noema 3.6+" platforms: ["iPhone", "iPad", "Mac", "Vision Pro"] reviewed: "July 26, 2026" canonical: "https://noemaai.com/docs/privacy-and-network-activity" --- # Privacy & Network Activity A feature-by-feature map of local, LAN, iCloud, Apple, Noema-operated, and third-party data paths. ## Local-first does not mean network-free Local models, documents, indexes, and conversations can remain on your device. Optional features can communicate with local-network, iCloud, Apple, Noema-operated, enterprise, or third-party services. This page describes those feature data paths; the Privacy Policy governs collection, retention, and legal terms. ## Feature-by-feature data flow | Feature | Destination | What may be sent | | --- | --- | --- | | Local models and RAG | Your device | Nothing leaves for generation or local retrieval. | | Local Python, Memory, Calculator, Converter, Charts | Your device | Tool inputs and results remain local; Python has no network access. | | Downloads | Selected model, dataset, or voice source | Asset identifier and normal connection metadata. | | Web Search | Selected search service and source reader | Query, locale, controls, candidate URLs or snippets, source references, and open/find parameters. | | Remote endpoint | Configured provider | Prompt, relevant chat and system context, supported attachments, tools and results, and retrieved excerpts. | | Autopilot remote router | Configured router | Each evaluated message, limited routing context, and capability information. | | Autopilot escalation | Configured stronger model | Prompt and context needed for the answer; knowledge excerpts only when explicitly allowed. | | Constellation chat sync | Private iCloud database | Chat and project payloads; models, datasets, and source files remain device-local. | | Constellation Local Network | Your Mac on the local network | Authenticated request, supported attachments, tool exchanges, and streamed response. | | Constellation Private Link | Your configured private mesh | Request, supported attachments, tool exchanges, and streamed response. | | Constellation Direct | The other personal device | End-to-end encrypted device-to-device request after CloudKit introduction. | | Noema Bridge | Noema-operated bridge | End-to-end encrypted traffic relayed as ciphertext; the Bridge does not run the model. | | Constellation Cloud Relay | Private iCloud database and your Mac | Request, optional bounded image assets, tool exchanges, and reply. | | Apple Speech | Apple speech-recognition system | Audio may be processed by Apple when on-device recognition is unavailable or not required. | | Local Whisper | Your device | Audio and transcripts remain local after download. | | Remote audio model | Configured provider | Media required for transcription. | | Calendar | Device Calendar store | Reads and confirmed writes use OS calendar permissions. | | Enterprise enrollment | Organization service | Enrollment credentials, policy state, catalogs, and managed dataset requests. | | Advertising attribution | Service named in the Privacy Policy | Limited install or campaign attribution described by that policy, excluding conversations and documents. | ## Off-grid Mode Off-grid Mode blocks external HTTP and HTTPS traffic initiated through Noema’s network stack. Local generation, local retrieval, and local tools continue to work. - Web Search, remote endpoints, Constellation sync and remote access, remote Autopilot routing and escalation, and new downloads stop. - Remote speech and audio services stop. - Local LAN and managed enterprise behavior can still be subject to platform or organization policy. - Off-grid Mode does not guarantee that separately launched software or operating-system services are isolated. ## Accounts, deletion, and support Personal local use does not require a Noema account. Constellation uses the Apple Account and private iCloud database already configured on the user’s devices. Third-party providers, Apple services, enterprise workspaces, and external tools may require accounts governed by their own terms. Deleting a local chat or dataset removes Noema’s local copy. Material previously sent to another provider remains subject to that provider’s retention policy. Review diagnostic exports before sending them to support. ## Related documentation - [Constellation Connections & Privacy](https://noemaai.com/docs/constellation-connections) - [Privacy FAQ](https://noemaai.com/docs/privacy-faq) - [About Noema](https://noemaai.com/docs/about-noema) - [Remote Endpoints](https://noemaai.com/docs/remote-endpoints) --- --- title: "Privacy FAQ" description: "Answers about offline use, optional network features, remote providers, Constellation, speech, enterprise services, and diagnostics." version: "Noema 3.6+" platforms: ["iPhone", "iPad", "Mac", "Vision Pro"] reviewed: "July 26, 2026" canonical: "https://noemaai.com/docs/privacy-faq" --- # Privacy FAQ Clear answers without absolute claims or fixed performance promises. ## Can Noema work without the internet? Yes. After compatible models and optional retrieval assets are installed, local chat, local RAG, and local tools can work offline. Network-backed features do not. ## What can leave my device? Only the features and routes you enable determine this. Web Search sends search and source-reading data; remote endpoints receive the context needed to answer; remote Autopilot can send routing or generation context; Constellation can sync through private iCloud and reach an awake Mac over Local Network, Private Link, Direct, Noema Bridge, or Cloud Relay; Apple Speech can use Apple services; enterprise and configured external tools follow their own disclosed paths. > See Privacy & Network Activity for the complete matrix. ## Does Noema guarantee Secure Enclave storage? No blanket guarantee applies to every model, file, database, provider credential, or platform. Noema uses platform storage and Keychain controls where the feature supports them; the operating system, backup configuration, and third-party services still define important boundaries. ## Are network services unlimited? Noema may provide features without an in-app quota or subscription, but external services can impose rate limits, outages, maintenance, account rules, or provider costs. ## How fast will a model run? There is no fixed tokens-per-second promise. Performance depends on device, thermal state, runtime, model architecture, quantization, context length, attachments, and settings. Use the built-in benchmark and runtime receipt on the device that matters. ## What should I check before sharing diagnostics? Diagnostics are designed to avoid chat and document contents and to redact common secrets, but you should still review every export before sending it. Third-party logs and screenshots may contain information outside Noema’s control. ## Related documentation - [Privacy & Network Activity](https://noemaai.com/docs/privacy-and-network-activity) - [About Noema](https://noemaai.com/docs/about-noema) - [Support](https://noemaai.com/docs/support) --- --- title: "Support" description: "Collect useful Noema diagnostics and contact support with the device, model, feature, and reproduction details needed to help." version: "Noema 3.6+" platforms: ["iPhone", "iPad", "Mac", "Vision Pro"] reviewed: "July 26, 2026" canonical: "https://noemaai.com/docs/support" --- # Support What to include in a report and where to find logs, diagnostics, and contact controls. ## Start with Diagnostics & Tools - Open Settings → Diagnostics & Tools. - Review device, storage, runtime, model, retrieval, and network checks relevant to the failure. - Use Share Logs when a reproducible issue needs technical detail. - Open Notes & Issues to record the behavior and Contact Support when you are ready to send it. ## Include these details - Noema version and build number. - Device model and operating-system version. - Model identifier, runtime format, and quantization. - Dataset, tool, remote endpoint, Constellation route, or Autopilot configuration involved. - Exact reproduction steps, expected result, and actual result. - A screenshot and timestamp when they help correlate logs. ## Contact support Use Contact Support in the app or email clientcare@noemaai.com. Do not send passwords, API keys, private model credentials, full sensitive documents, or unreviewed diagnostic exports. Response time depends on issue volume and complexity; there is no published one-business-day SLA. ## Before sending - Confirm the issue still occurs on the latest public Noema build. - Try a smaller compatible model or simpler route when the failure appears memory-related. - Check Off-grid Mode when a network-backed feature unexpectedly stops. - Review the relevant documentation page and remove secrets from screenshots and logs. ## Related documentation - [Constellation Model Ejection & Help](https://noemaai.com/docs/constellation-troubleshooting) - [Quick Start](https://noemaai.com/docs/quick-start) - [Model Settings](https://noemaai.com/docs/model-settings) - [Privacy FAQ](https://noemaai.com/docs/privacy-faq)