AI ORCHESTRATION PLATFORM

Platform Architecture

A deep dive into the orchestration layer, state architecture, precision editing pipelines, and multi-tool infrastructure powering the TwelveAgents operating environment.

Introduction

TwelveAgents is a proprietary, media-rich orchestration environment that brings twelve specialized AI agents together inside a single, unified operating system for AI work. It exists to solve a problem that a raw chat interface never can: turning open-ended reasoning into finished, reliable, multi-step output — documents, websites, videos, decks, code, and research — without the user having to manage a single intermediate step by hand.

The architecture is built around five pillars. Together, they form the operating environment inside which reasoning becomes completed work:

  • A unified agent runtime that gives every specialized agent a consistent, durable operating envelope, regardless of the underlying reasoning engine.
  • A deterministic state and consistency engine that structurally prevents the contradiction, drift, and hallucination failure modes inherent to conversational AI.
  • A surgical editing and version-control engine that treats every generated artifact as a precisely maintainable asset rather than something to be blindly rewritten.
  • A native, cross-format media pipeline that lets agents chain image, video, audio, and document operations together in a single request without losing fidelity.
  • An on-device privacy layer that gives users direct, local control over what sensitive information ever leaves their device.

The Unified Agent Runtime

TwelveAgents ships with twelve specialized agents, each with a distinct persona and area of focus, but all twelve run inside the same governing runtime shell. This orchestration shell — not any single LLM — is what defines an "agent" in the platform's architecture. It is the shell that carries identity, tool access, context, execution capability, and state forward, turn after turn, reducing the reasoning engine to a swappable component.

The Four Layers of the Runtime Shell

Persona & Instruction Layer

Every agent carries a durable identity anchored to the runtime itself. This persists unchanged through model swaps, context compression events, and multi-day gaps between sessions.

Tool Accessibility Layer

Every agent has latent access to the platform's full toolkit. If a request benefits from a disabled capability, the runtime detects the gap and prompts you to extend the toolset mid-conversation.

Context Management Layer

Context is not just "everything said so far." TwelveAgents exposes user-controlled context injection, allowing you to tune continuity, relevance, and resource cost per turn.

Execution & State Layer

Every agent operates against a shared, persistent execution substrate that survives individual reasoning calls. It remembers not just what was said, but what was built, decided, and computed.

Adaptive Tool Extensibility

Because tool access is a property of the orchestration runtime and not of any individual agent persona, the platform offers adaptive extensibility. Any agent can reach for any tool in the full toolkit the moment a task calls for it. A request never dead-ends just because the currently active agent's default configuration didn't anticipate it — the runtime recognizes the gap and offers to close it, live, mid-conversation.

Multi-Chat Independence

Every chat is an independent instance of this runtime: its own agent, its own history, its own context settings, its own enabled toolset. Users can run as many parallel chats as they need, completely isolated from one another, without cross-contamination of state or files.

Meet the Twelve Agents

Agent Focus Description
Nexo · AdaptiveUniversal assistantDiscusses any topic and automatically identifies and suggests the right tools to enable as a task develops.
Zeta · Quick ChatEveryday companionHandles everyday questions, quick math, live web search, and calendar events with minimal overhead.
Lexa · Writing StudioCopywritingDrafts essays, blogs, short stories, and high-converting creative copy.
Doca · Document StudioDocument specialistReads, writes, and exports CVs, reports, technical documentation, and architecture diagrams.
Anix · Research & AnalysisLead researcherConducts deep web research, analyzes complex documents, and exports findings into polished presentations.
Dato · Data & FinanceData analysisAnalyzes spreadsheets, performs financial calculations, manages structured data, and visualizes it into charts.
Sino · Creative StudioCreative directionGenerates media from scratch — images, logos, video, and music — then stitches, dubs, subtitles, and edits.
Pixo · Media EditorPost-productionTakes uploaded files and performs stitching, cropping, trimming, subtitling, translation, dubbing, filtering, and vectorization.
Vixo · Vision & MediaMedia analysisExtracts text, summarizes content, and surfaces insights from uploaded images, video, and audio.
Sona · Social MediaSocial contentCreates and formats posts, reels, and captions sized correctly for each platform.
Kado · Code StudioPair programmingWrites, executes, and debugs code, and searches the web for API documentation as needed.
Maku · Website BuilderWeb developmentDesigns and generates fully coded, responsive HTML/CSS website templates ready to download.

Deterministic Multi-Turn State Architecture

The Problem: Conversational reasoning is stateless and probabilistic. Ask an unmanaged AI system to recompute a number it already gave you, and there is no structural guarantee it returns the same value twice. Ask it to modify a document, and it will often reconstruct the whole thing from imperfect memory—introducing drift in wording, structure, or formatting.

TwelveAgents closes this gap with a dedicated state and consistency orchestration engine that sits above every individual reasoning and tool call, enforcing guarantees that no single LLM call is structurally capable of making alone:

  • Canonical Fact Locking: The first time a specific number, brand decision, or foundational value is established, it is written into the conversation's canonical state and locked. Any later attempt by any tool to introduce a conflicting value is intercepted and corrected automatically.
  • Deterministic Computation: Arithmetic that matters (totals, growth rates) is never left to the reasoning engine's probabilistic math. Calculations are executed through dedicated TypeScript routines, producing exact, reproducible results saved to the ledger.
  • Pre-Render Validation: Before any document or presentation compiles, its internal figures are checked against the canonical values locked in state. If a breakdown does not reconcile with its stated total, generation halts and the workflow is redirected to correct the error before anything is shown to the user.
  • Cross-Turn Artifact Memory: Every generated file has its exact structural composition retained in state. An edit requested days later recalls the artifact's true prior structure, applying only the requested change while leaving untouched elements identical.
  • Brand & Aesthetic Continuity: Once a color palette or typography pairing is established, every subsequent generated asset (logo, slide deck, webpage) inherits it automatically.

The Combined Effect

Individually, each mechanism closes a specific failure mode. In combination, they produce a system that is architecturally prevented from silently contradicting itself, forgetting a prior decision, or letting numbers drift across a long, multi-step piece of work — regardless of how many tools, turns, or assets are involved.

Surgical Editing & Artifact Version Control

When a task involves a long-form document, a multi-page website, or a codebase, the naive approach — regenerating the entire artifact from scratch on every edit request — is slow, expensive, and structurally risky.

TwelveAgents never regenerates blindly. Its editing architecture works the way disciplined engineers work with version-controlled systems: the orchestration layer identifies the smallest precise change required, applies it directly against the artifact's true underlying structure, and deterministically recompiles the asset.

This is visible directly in the product experience. The platform produces a structured, line-level account of exactly what was altered:

📍 CHANGE 1 (Line ~15): - return MaterialApp( - title: 'Complex Number Adder', + return MaterialApp( + title: 'Complex Number Adder by Kado',
📍 CHANGE 2 (Line ~152): - appBar: AppBar( - title: const Text('Complex Number Adder'), + appBar: AppBar( + title: const Text('Complex Number Adder by Kado'), Summary: +4 lines added, -4 lines removed (title and AppBar title). Everything else remains identical.

This provides unparalleled Speed (only the affected region is touched), Safety (untouched sections are verifiably identical), Auditability (every version is retained and restorable), and Transparency (the user sees a clear diff, never needing to guess what changed).

Full Lifecycle Coverage

Stage What Happens
CreationThe full asset is generated from scratch, and its exact structure is committed to state as the artifact's baseline version.
Iterative editingEach subsequent change is applied as a targeted modification against the last known-good structure — never a blind rewrite.
Version comparisonAny two points in an artifact's history can be compared directly, producing a precise, human-readable account of what changed.
RollbackReverting to any retained prior version (the last N versions per artifact) is an exact, deterministic restoration of that version's true saved structure, not an imperfect regeneration attempt.

Developer Experience: Code Generation & Rendering

Code as a First-Class Managed Artifact

Generated code is treated with the same rigorous version control. Every script, component, or multi-file project is versioned, diffable, and edited surgically. A working session can span many turns, with each request ("rename this variable," "add error handling") applied as an auditable change against the project's real current state.

"I've got solid control over this: I can generate and edit code seamlessly, track artifacts with version history, apply surgical patches to existing files, and diff versions to show exactly what changed."
— Agent Kado

Full Display & Rendering Customization

Because code and mathematical notation are rendered directly inside the conversation, TwelveAgents gives users granular control over exactly how that output looks:

  • Code Theme: Choose from dozens of syntax-highlighting themes (from high-contrast to editor-inspired palettes), previewed live.
  • Line Numbers & Dividers: Toggle numbered lines and visual dividers for dense code readability.
  • LaTeX Math: Render mathematical notation properly using LaTeX typesetting so equations read the way they would in a textbook.
  • Math Buffering: Buffer mathematical rendering briefly to prevent visual flicker as complex equations stream in.

TwelveAgents treats every generated line of code as a managed artifact — versioned, diffable, and built around a full, customizable authoring environment.

Native Media Intelligence

TwelveAgents treats every image, video clip, audio file, and document as a first-class object that the platform itself owns, tracks, and manages. Media moves through the system natively: an agent can generate, transform, analyze, and chain operations together without losing fidelity or requiring the user to manually re-upload intermediate files.

What a Single Request Can Chain Together

This architecture allows deep, multi-tool workflows to execute from a single natural-language instruction. The orchestration layer maps the path; the tools execute the steps.

Simple, single-step:
  • Generate a square image of a cute astronaut cat.
  • Convert this HTML into clean Markdown.
A few steps chained:
  • Generate a square image of a cute astronaut cat, then make it black and white.
  • Find today's top AI headline, summarize it, and read it back to me in a British female voice.
Deeper chains, several tools working in sequence:
  • Search the web for how solar panels work, generate a diagram explaining it, add a narrated audio track, and build it all into a downloadable website.
  • Generate a video of a waterfall, speed it up 2x, resize it to a square for Instagram, then compress it so it's small enough to email.
Full multi-tool workflows:
  • Calculate the number of days until 2030, search for predicted future technologies, diagram those technologies, generate a video of a futuristic city, write and record a voiceover summarizing it all, merge the voiceover onto the video, and package everything — diagram, video, and text — into a downloadable website.

Worked Example: The Mars Hero Video

To illustrate how deep a single creative chain can go natively, consider the platform's own 22-second hero video production, built entirely through sequential instructions to Agent Sino, with no external editing software:

1. Base Generation
An uploaded mascot image is animated into a cinematic wide shot with a matched score and ambient audio.
2. Native Continuity Extensions
The clip is extended forward in time across two separate requests while explicitly preserving character design continuity and creating a synchronized lighting climax.
3. Audio Generation & Mixing
A separate vocal track is generated and automatically blended underneath the video's existing native sound effects, ducked to the correct level.
4. Precision Editing & Finalization
A blank opening frame is identified and trimmed natively. A specific frame is pulled out at 16s to serve as a thumbnail, and automatically embedded as the video's metadata.

Media as a Managed Asset Class

Every media asset is securely isolated per account, maintains referential integrity (it can be re-referenced across many turns without re-uploading), supports cross-format conversions (like SVG vectorization), and follows a strict managed retention lifecycle to keep the system fast.

Reliability, Governance & Execution Limits

A platform capable of chaining a dozen tools together inside a single request must structurally guarantee that it cannot run away, stall indefinitely, or fail silently. TwelveAgents treats reliability as a property of the architecture itself, enforcing explicit ceilings.

Capability General Limit (per request)
Video generation / dubbing1
Music generation4
Image generation20
Audio generation / processing4
Code executions6
Deep media analysis / transcription12
Media editing operations8
Document / file generation16
Web searches8
Lightweight utilities120
Agent workflow iterations16
Maximum workflow runtime15 minutes
  • Graceful Partial Completion: If a request exceeds a limit partway through, the platform does not fail the whole request. It completes what it safely can, explains exactly what limit was hit, and offers a concrete path to continue.
  • Isolated Sandboxed Execution: Any code generated and executed runs inside a fully isolated, ephemeral sandbox environment, structurally separated from the platform's core infrastructure.
  • Verified Output: Numeric and structural correctness are checked against canonical state before a finished asset is delivered. Errors are caught upstream by the orchestration layer, not downstream by the user.

Storage & Account Architecture

Every account operates inside a securely isolated storage boundary, with usage tracked continuously and enforced in real time.

Resource Allowance
Cloud storage1 GB per account
Active chat roomsUp to 60 per account
Generated media retention7 days from date of generation
  • Managed Media Lifecycle: AI-generated media is retained for a fixed 7-day active window. If an account hits 1 GB before the window elapses, the platform automatically reclaims media from the oldest chats first, leaving the text conversation intact.
  • Real-Time Quota Enforcement: Consumption is measured at the moment of use. A write that exceeds capacity is rejected cleanly before it happens.
  • Verified Deletion: Deleting a conversation triggers a genuine, thorough removal of associated assets—not a soft flag that leaves data recoverable.
  • Continuous Self-Auditing: The platform continuously reconciles storage records against physical usage in the background, preventing quota drift over time.

On-Device Privacy Layer: Prompt Protection

TwelveAgents includes an optional, built-in privacy safeguard called Prompt Protection. It is an on-device sensitive-data detection and redaction layer that scans content before it is ever sent, giving users a direct, local control point over what leaves their device.

No content is uploaded to a remote service to perform this scan — the detection runs entirely locally using a mature, signature-based contextual pattern-matching engine.

Protection Level Detects & Redacts Locally
Secrets API keys, JWTs, session IDs, contextual passwords, database connection strings, cryptographic private keys.
Balanced Everything in Secrets, plus email addresses, phone numbers, contextual IDs, US SSNs, device IDs (IMEI/MEID), GPS coordinates, and biometric data references.
Standard Everything in Balanced, plus credit card numbers, IBANs, SWIFT/BIC codes, and cryptocurrency addresses.
Enhanced Everything in Standard, plus IP addresses, MAC addresses, UUIDs, URLs, file paths, and vehicle identification numbers (VINs).

An Honest Accounting of Its Limits

Prompt Protection is pattern-based plain-text scanning, not a magic guarantee. It may miss sensitive text in unusual formats, and it does not scan inside binary files like PDFs. It is a powerful local safeguard, not a replacement for regulatory compliance or human review.

Agent Brains & Transparent Credit Model

The Orchestration Layer dictates identity, manages toolchains, handles surgical editing, enforces constraints, and persists state. The underlying reasoning step that powers it—the "Agent Brain"—is simply a pluggable component executed via curated AI providers (LLMs).

The choice of Agent Brain can be changed by the user at any time, including mid-conversation. Because the platform's orchestration remains constant, the agent's identity, locked facts, enabled tools, and multi-turn artifact history all carry forward completely unaffected when the brain is swapped.

Transparent Credit System

TwelveAgents uses a simple credit system with no subscriptions and no auto-renewals. Credit packs are purchased when needed and used at your own pace. Purchases are handled natively through your app store (Apple, Google, Microsoft) and synchronize automatically across all your supported devices.

Pack Price
Starter$5
Standard$19
Plus$49
Pro$99
Studio$499

Closing Summary

The orchestration layer described across this document — the unified agent runtime, the deterministic state engine, the surgical version-control architecture, the native media pipeline, the execution governance, and the on-device privacy layer — is what turns a raw reasoning engine into a platform capable of finishing real, production-grade work.

Each layer exists to close a specific failure mode of conversational AI: contradiction, drift, unsafe regeneration, dropped context, runaway execution, and uncontrolled data exposure. Together, they form the operating environment inside which twelve specialized agents function as one coherent, dependable system.