ODINO 3.11 · NEMOTRON 3.5 LIGHTNING · PRODUCTION LIVE
Thirty-billion-parameter capacity.Sustained at the sovereign edge.
Odino 3.11 promotes NVIDIA Nemotron 3.5 Lightning 30B-A3B NVFP4 to production on a single NVIDIA Jetson AGX Thor. The 30B Mixture-of-Experts model activates approximately 3B parameters per token: 7.5× the total capacity of the former Nano-4B model, with 72.623 tok/s sustained production throughput and 68.072 ms median time to first content.
DIE EINSCHRÄNKUNG
More capability, inside the same sovereign machine.
Jetson AGX Thor 5000 provides 128 GiB of unified memory — substantial for an edge node, finite for a complete multimodal cognitive system. Odino began as the answer to that hard resource boundary. Its current production engine now holds a 30B MoE model at 19,102 MiB GPU UMA residency while Nemotron Omni, Riva speech and the governed service fleet remain live.
The advance is not a like-for-like model swap. Nemotron 3.5 Lightning brings 7.5× greater total parameter capacity than Nano-4B while activating about 3B parameters per token. This preserves a relatively contained active-compute footprint and materially expands the system's scientific reasoning capacity.
Q19 · PRODUCTION PERFORMANCE · 23 AUGUST 2026
+28% sustained. +36% qualified. Up to +63% native.
The headline uses the last sustained, comparable Odino Nano result as its baseline. All measurements run on the same Jetson AGX Thor; qualified HTTP and native results use the fixed Q18 workload, while the production figure is the mean of five HTTP runs.
72.623
TOK/S · PRODUCTION
77.104
TOK/S · HTTP QUALIFIED
92.8
TOK/S · NATIVE PEAK
68.072 ms
MEDIAN FIRST CONTENT
| Arbeitsbelastung | Comparable Nano baseline | Odino 3.11 | Improvement |
|---|---|---|---|
| Sustained production · five-run HTTP | 56.895 tok/s | 72.623 tok/s | +27.6% |
| Qualified HTTP · fixed Q18 workload | 56.895 tok/s | 77.104 tok/s | +35.5% |
| Native engine · fixed Q18 workload | 56.895 tok/s | 92.8 tok/s | +63.1% |
| Time to first content · production median | — | 68.072 ms | CURRENT |
Method boundary: an older Q1 run recorded 97.651 tok/s on a short, mixed workload. It is not used as a direct comparator because prompt, model, runtime and methodology differ.
SYSTEM OUTCOME · SAME Q9 SCIENTIFIC MANIFEST
The larger result is capability and scientific reliability.
On the fixed bilingual scientific and adversarial benchmark, Odino now closes every measured case. Deterministic Direct Evidence handles exact graph, process and numeric questions with provenance gates and zero LLM calls.
| Metrisch | Q9 baseline | Current system | Veränderung |
|---|---|---|---|
| Scientific accuracy · fixed Q9 manifest | 63.33% | 100% | +36.67 points · +57.9% relative |
| Direct Evidence latency · p50 | 3,556.867 ms | 31.370 ms | −99.1% · 113.38× faster |
| Grounded catalogue request | HTTP 413 | HTTP 200 · 5.73 s | Real tool trace · 236 rows |
Knowledge Graph, claim-level provenance, evidence assessment, local UQ/OOD, grounded tools and specialist read-only consumers now form a single governed evidence path. The production graph contains 259 nodes and 246 edges.
These capabilities were added without enabling mutation, export, ingest or autonomous experiment execution. Workbench and vector-store invariants, Guardian policy, kill switch and human-authority boundaries remain intact.
HISTORICAL FOUNDATION · PHASE 1
Die Zahlen, die den Bau rechtfertigten.
Die folgende Tabelle vergleicht Odino Phase 1 mit der vorherigen vLLM-Konfiguration auf derselben Jetson AGX Thor 5000-Hardware, die dasselbe FP8-quantisierte Chat-Modell ausführt. Validiert am 06.12.2025.
| Metrisch | vLLM (vorher) | Odino Phase 1 (validiert) | Δ vs. vLLM | Odino Sprint 3-Streaming |
|---|---|---|---|---|
| Speicherbelegung (untätig) | ~67 GiB | 381 MiB | −99.4% | ~381 MiB (unverändert) |
| Speicherbelegung (geladen) | ~70 GiB (gesättigt) | ~15 GiB | −78% | ~15 GiB (unverändert) |
| Kaltstartzeit | 30–60 s | < 2 s | −96% | < 2 s Motor + erstes Aufwärmen |
| TTFT streamen | n/a | n/a | n/a | warmer stationärer Zustand innerhalb des Sub-200-ms-Bandes |
| Interaktive Unterstützung | n/a | Nur synchron | n/a | Token-für-Token-SSE |
Der wiederhergestellte Speicher – etwa 50 GiB, der durch den Vergleich des geladenen Zustands frei wird – ermöglicht es dem Cluster Tin Man, Vision- (RADIO ViT-H/16), Audio- (Parakeet Conformer) und Echtzeit-Wahrnehmungskerne neben dem Argumentationskern auf einem einzelnen Jetson Thor-Knoten zu betreiben.
TECHNISCHE GRUNDSÄTZE
Native compilation. Bounded residency. Governed operation.
Drei architektonische Entscheidungen, einheitlich umgesetzt:
- Thor-native compilationA TensorRT Edge-LLM 0.10.0 engine compiled for Jetson AGX Thor runs Nemotron 3.5 Lightning in NVFP4, without a general-purpose serving framework in the inference path.
- Bounded residencyThe production engine holds full residency at 19,102 MiB of GPU unified memory while the wider multimodal stack remains operational.
- Governed production pathThe canonical model endpoint is fixed, the former Nano engine is inactive, mutating tools remain disabled, and supervised readiness is tested across the complete stack.
Each decision reflects Odino's role as a purpose-built reasoning runtime for Tin Man on this specific hardware, rather than a general-purpose inference framework.
DETERMINISM
Evidence paths before generative paths.
Beyond raw throughput, Odino routes exact graph, numeric and process questions through a deterministic Direct Evidence path. Those answers carry bounded local evidence and provenance, bypass generation, and require zero LLM calls. Ambiguous reasoning remains on the governed model path rather than being misrepresented as deterministic evidence.
The architecture keeps the distinction visible: evidence-backed facts, model reasoning and tool results retain separate provenance. Shield Brain and the human-authority boundary apply across every route, while mutating tools and autonomous experiment execution stay disabled.
INTEGRATION
Eine Komponente, kein Produkt.
Odino wird nicht als eigenständiges Produkt geliefert. Es ist der Chat Core des kognitiven Clusters Tin Man – die Argumentationsschicht, die Eingaben im Textmodus interpretiert und die multimodale Eingabeaufforderungsassembly erstellt, die nachgeschaltete Kerne nutzen. Sein Wert wird durch die Integration mit Vision-, Audio-, Brainstem-, Speicher- und Echtzeitkernen unter der Shield Brain-Steuerungsarchitektur realisiert.
The 3.11 production line exposes one canonical Nemotron 3.5 endpoint with no automatic model fallback. Its read-only scientific consumers, grounded tools and deterministic evidence routes share the same Guardian and human-authority boundary; mutating operations remain disabled.
SPRINT 3-VERBESSERUNG · MAI 2026
Streaming-Token-für-Token-Zustellung.
Im Mai 2026 wurde die Odino-Laufzeit um serverseitige Streaming-Unterstützung erweitert, wodurch der Tin Man-Cluster die Generierung von Sprachmodell-Tokens mit nachgelagerten Syntheseschichten (Text-to-Speech) kanalisieren und die Latenz der End-to-End-Sprachschleife deutlich unter 1 Sekunde reduzieren kann.
Die Abwärtskompatibilität mit dem synchronen Endpunkt der Phase 1 bleibt erhalten. Streaming wird über OpenAI-kompatible Server-Sent Events implementiert und liefert eine inkrementelle Ausgabe pro Token.
Empirische Leistung:
- Warmer Dauerzustand bis zum ersten Token: deutlich innerhalb des Zielbands unter 200 ms
- Kohärente Reaktion auf das Drohnen-Kommandokorpus auf interaktivem Niveau
Der Golden-Master-Motorplan der Phase 1 bleibt unverändert erhalten; Das Streaming-Verhalten wird durch eine unterschiedliche Servernutzung erreicht, die zur Laufzeit denselben Engine-Plan bindet – ein Herkunftsmuster, das darauf ausgelegt ist, validierte Assets über iterative Verbesserungen hinweg stabil zu halten.
ENGAGE
Technische Tieftauchgänge werden im Rahmen einer Partnerschaft veröffentlicht.
The full Odino technical report, benchmark manifests, integration notes for the Tin Man cluster, engine-build evidence, and the embedded-AI engineering pathway are released under non-disclosure agreement. Engage to begin a conversation.
Direkt: engage@reinventy-solutions.ca