SABHAin sessionadvisors 5 groundeddecisions filed 46survived prosecution 100%awaiting your tap 1 gate
arivu · the faculty of judgment
arivuIt decides. Then it acts.
arivu · Tamil — intellect, wisdom — the determining mind; the faculty that resolves doubt into settled judgment.
A chamber of grounded advisors that deliberate, survive their own prosecution, and on one approval execute the decision.
arivu is a company's DECIDE faculty made operational. It does not answer questions of fact or essence — that is manas, the memory; arivu is convened only when something is at stake. Independent expert lenses argue from the org's own live numbers; a chair reconciles them into one verdict; an adversary tries to kill that verdict before it ships; a single human signs off; and then real execution fires across every team — all on the record. It commits a live change and dispatches A2A commands — and the world hears it only through kural, the company's only mouth. It commits and commands; it never merely recommends.
Submitted to the Google for Startups AI Agents Challenge — Track 1 · Build. A multi-agent system built on the Agent Development Kit (ADK), powered by Gemini for the chamber's chair and five advisors, and Claude via Vertex AI for the two highest-stakes reasoning steps: the verdict synthesis and the adversarial prosecution.
RUN
◈ The deliberation, live · watch doubt crystallize
Noise → crystal → executed reality.
The founder's question sits at the rim of the chamber. Watch the five mantris fan out, argue from their own grounded numbers, the debate loop snap the noise into a crystallizing lattice — and the Risk mantri catch the churn cliff a lone analyst skips. The Claude·Vertex chair strikes the verdict; the Claude·Vertex Prosecutor fires its counter-force and the seal nearly shatters — until defensibility crosses 0.8 and it re-forms harder. Then one APPROVE commits the change, dispatches the orders, and surfaces the published resolution. Flip to lone analyst to see the thin, pale verdict that misses the cliff.
# arivu · chamber transcriptframe» decompose "raise Pro to $39?" · ground via tools: manas A2A + admin_stats + admin_analyticsparallel» 5 mantris argue from disjoint data — positions form before any sees another'srisk»CATCHES the churn cliff at $39 — the lone path misses itchair»[Claude · Vertex] verdict: Raise to $34, grandfather existing, 30-day noticeprosecutor»[Claude · Vertex] steelmans do-nothing · loop until defensibility ≥ 0.80gate» seal re-forms · conf 0.88 · dissent noted · HALT — awaiting founder approval# press Run (JS) to watch doubt crystallize · then Approve to execute
01
The business problem · who it's for
A loaded call, and no board to put it to.
Sundara Coffee Co. — a four-person DTC coffee brand — has to make a real strategic call this quarter: "Should we raise our Pro subscription to $39?" A real company would put this to a board, a fractional CFO, a growth lead, and a brand head, fight it out, pressure-test the downside, decide, and then make every team act on the decision in lockstep. Sundara has none of those people and can't afford a single one of those hires.
So the call either gets made on a founder's gut at 1am — or, worse, never gets made, and the question rots. The gap isn't "give me an opinion." Any chatbot does that, badly, from model memory. The gap is the whole apparatus: independent expert lenses arguing from the company's own live numbers, a chair that reconciles them into one verdict, an adversary that tries to kill that verdict before it ships, a single human sign-off, and then real execution across every team — all on the record.
The missing board
No CFO, no growth lead, no brand head
A four-person company has no one to argue the unit economics against the positioning against the downside. The lenses that would catch each other's blind spots simply don't exist.
The 1am gut call
A loaded decision made alone, unrecorded
With no one to deliberate with, the founder decides from instinct at midnight — or defers it indefinitely. Either way the reasoning is never written down, and the dissent is never heard.
The chatbot trap
An opinion from memory is not a decision
Asking a model "what should I do?" returns a confident paragraph from training data — ungrounded in the company's actual revenue, churn, and brand promises. It recommends; it cannot commit or act.
The execution gap
A decision no one carries out is no decision
Even a good call dies if the teams don't move in lockstep. The price changes but the campaign never launches, the banner never ships, the config never flips. A verdict has to become action.
4
people at Sundara — and no board
$0
budget for a fractional CFO or growth lead
1am
when the gut call tends to get made
1 tap
what arivu asks of you before it acts
arivu is the company's DECIDE faculty made operational. Where manas gathers, registers, and doubts, arivu resolves doubt into settled judgment — and acts on it. It is the board, the CFO, the growth lead, and the brand head you couldn't hire, convened in one chamber, grounded in your own numbers, and bound to a single approval.
02
What it does · end to end
From a loaded question to a published resolution.
A strategic question enters the chamber. The chair frames it and grounds it in the org's own live data. Five disjoint lenses argue, the debate loop converges, Claude synthesizes a verdict, Claude prosecutes it — and the pipeline halts at one gate. On approval, arivu commits the change, dispatches the orders, and files the resolution.
→ 01
The chair frames & grounds · Gemini 2.5 Pro · TOOLS, not agents
The coordinator decomposes the founder's loaded question into the sub-questions each lens must answer, and pulls grounding via tools: manas over A2A (brand, voice, strategy) and the opt-in example MCP tools admin_stats and admin_analytics (enabled once the transport is confirmed) — so every advisor argues from the org's own numbers, never model memory. A single MCP call is a tool, not an agent.
→ 02
Five mantris deliberate in parallel · Gemini 2.5 Flash · ParallelAgent
The Economist, Growth, Brand, Risk, and Ops lenses fan out simultaneously — positions form before any advisor sees another's, so the chamber is anti-groupthink by construction. Each has a disjoint lens and disjoint data source, so they cannot collapse into one agent. This is where the Risk mantri catches the churn cliff the lone path misses.
→ 03
The debate loop converges · Gemini · LoopAgent
Advisors rebut each other's grounded claims over N critique/rebuttal rounds. The loop terminates on a deterministic numeric threshold — a converged-confidence value or a max-round bound — never on "the advisors agreed." Bounded, safe by construction, no runaway deliberation.
→ 04
The chair strikes the verdict · Claude via Vertex AI
Synthesis-under-conflict is the highest-stakes reasoning step, so it runs on Claude. It reconciles the surviving positions into one verdict — a specific, executable decision with cited reasons, the preserved dissent, and a numeric confidence — as structured JSON the executor can act on.
→ 05
The Prosecutor tries to defeat it · Claude via Vertex AI · LoopAgent
The grafted self-verification gate. The Prosecutor must steelman the do-nothing case and try to defeat the verdict on the merits. A loop re-runs verdict ↔ prosecution until the verdict survives at defensibility ≥ 0.80 (tool_context.actions.escalate fires on threshold) — or a deterministic max-iter rollback returns "no safe decision — re-frame." The chamber prosecutes itself before it ever asks the human to trust it.
◈ 06
One gate — the pipeline halts
Exactly one human approval stands between deliberation and irreversible action. Durable, resumable run state (the orchestrator's run registry; Vertex Memory Bank is roadmap) holds the chamber transcript, so an approval that lands later resumes where it paused. On APPROVE → the executor fires. On REJECT → the dissent is captured and the chamber re-enters the loop.
▣ 07
The executor acts — atomically, on the record
arivu commits the live change the verdict implies — a feature-flag flip or campaign-budget reallocation (a config commit, deliberately never a price/revenue column write). It dispatches A2A commands to the two executors — kural launches the announcement campaign, kalai renders the banner (real token spend). The dispatch is ×2, but the banner reaches the world only through kural, the company's only mouth: kalai hands its approved master to kural, and kural formats and publishes it per channel only behind the founder's publish sign-off. And it files a signed board resolution via docs_create() then docs_publish() to a real URL, plus a planner_create() entry to track follow-through. The artifact is the published resolution; the irreversible action is the committed change plus the dispatch that triggers real spend and a live campaign carried out at the mouth.
arivu · chamber transcript
# 01:14:22 question receivedchair» frame "raise Pro to $39?" → ground: manas A2A + admin_stats + admin_analyticseconomist» contribution margin holds; $39 leaves margin on the table only if churn stays flatgrowth» +price reads as a positioning signal; warns of conversion drag on top-of-funnelbrand» manas canon: grandfathering existing users is a brand-trust requirement, not a nicetyrisk» activity-analytics: churn cliff at $39 — cohort retention breaks past $36ops» deploy + billing signals: flag flip is low blast-radius; safe to ship nowchair»[Claude · Vertex] verdict: Raise to $34, grandfather existing, 30-day noticeprosecutor»[Claude · Vertex] steelman do-nothing → re-verdict → defensibility 0.84 ≥ 0.80gate» conf 0.88 · dissent (Growth) noted · HALT — awaiting founder approvalapproved»flag flipped · A2A → kural/kalai · resolution published
03
The multi-agent org · the sabha
Five lenses. A chair. An adversary.
arivu is a real chamber, not a single prompt. Each advisor is an ADK LlmAgent with its own model, instruction, and tools. The five mantris are genuine agents because their lenses and data sources are disjoint and cannot collapse into one. Routine deliberation runs on Gemini. The two highest-stakes reasoning steps — synthesizing the verdict and prosecuting it — run on Claude via Vertex AI, precisely because reasoning under conflict is where a stronger model earns its place.
Gemini — the chair & the five grounded mantris Claude via Vertex AI — synthesis & prosecution (highest stakes)
arivucoordinator · chair of the sabha · root_agent
The chamber chair and orchestrator. Frames the founder's loaded question into sub-questions and grounds them via tools — manas over A2A plus the live example MCP calls admin_stats, admin_analytics — then runs the convergence pipeline and halts at the single HITL gate. On approval it fires the executor: commits the change, dispatches A2A commands, files the resolution. Orchestrator only — it does not vote.
A genuine agent — disjoint lens, disjoint data. Owns the money math: LTV, CAC, contribution margin, the elasticity of a price move. Tools: admin_stats (revenue, paying users) + a margin/elasticity calculator. In the demo it argues a $39 list leaves margin on the table only if churn holds.
Model · Gemini 2.5 Flashdata · admin_stats + elasticity tool
Growth advocate mantrifunnel & acquisition lens
A genuine agent. Argues the top-of-funnel case from kural's pipeline data over A2A plus admin_analytics(type=user-growth). In the demo it pushes for a higher price as a positioning signal — and warns about conversion drag. Its position becomes the preserved dissent in the final resolution.
Model · Gemini 2.5 Flashdata · admin_analytics + kural A2A
A genuine agent grounded in manas over A2A — brand canon, voice, positioning, prior commitments to customers. Its job is to flag whether a move is on-brand and whether it breaks an implicit promise: e.g. that grandfathering existing users is a brand-trust requirement, not a nicety.
Model · Gemini 2.5 Flashdata · manas A2A retrieval
Risk · devil's-advocate mantridownside-first lens
A genuine agent whose lens is downside-first. Tools: admin_analytics(type=activity) for churn/cohort signals + a scenario-stress tool. In the demo it surfaces the churn cliff the lone-analyst path misses — the single sharpest reason the full council beats one model answering alone.
Model · Gemini 2.5 Flashdata · activity-analytics + stress tool
Ops-feasibility mantrican-we-ship-this lens
A genuine agent that checks whether the company can actually ship and operate a decision, against live ops signals — deploy health, config-change risk, billing-system blast radius. Vetoes decisions that are sound on paper but unsafe to execute now.
Model · Gemini 2.5 Flashdata · live ops-signal query
Chair-synthesizerverdict author
The high-stakes DECIDE agent. Reconciles the surviving advisor positions after the debate loop into one verdict — a specific, executable decision with cited reasons, the preserved dissent, and a numeric confidence. Claude is chosen here precisely because synthesis-under-conflict is the highest-stakes reasoning step in the pipeline. Outputs structured verdict JSON the executor can act on.
Model · Claude (Opus-class) — Vertex AIpattern · AgentTool · structured verdict JSON
Prosecutoradversarial null-case
The grafted self-verification gate. After a verdict is struck, the Prosecutor must steelman the do-nothing case and try to defeat the verdict on the merits. A LoopAgent re-runs verdict ↔ prosecution until the verdict survives at defensibility ≥ 0.80 (tool_context.actions.escalate on threshold) — or a deterministic max-iter rollback returns "no safe decision — re-frame." Claude again, because beating a strong argument is itself a high-stakes reasoning task. This is the seal nearly shattering, then re-forming harder.
Model · Claude (Opus-class) — Vertex AIpattern · LoopAgent · escalate @ 0.80 + max-iter rollback
The convergence — why a chamber beats one analyst
5 mantris
Gemini · parallel · disjoint data
→
Debate loop
numeric threshold exit
→
Chair · verdict
Claude · Vertex · synthesis
→
Prosecutor
Claude · ≥ 0.80 or re-frame
→
One gate → act
commit · dispatch · publish
A single model answering "what should I do?" judges its own opinion from memory and never surfaces the dissent — so it ships the $39 churn cliff it never saw. arivu splits the labor: five disjoint lenses argue from the org's own numbers, a deterministic debate loop converges them, a separate Claude chair reconciles them into one verdict with preserved dissent, and a Claude Prosecutor must defeat that verdict before it survives at defensibility ≥ 0.80. The division between advisor, chair, and adversary — and the deterministic thresholds between every loop — is the safety property that makes the chamber measurably better than one model deciding alone.
04
How it's built · the ADK bar
The one place Parallel and Loop are earned.
arivu is built to the Google-grade bar, not bolted together. It is scaffolded with the agent-starter-pack, ships a checked-in eval set graded at the @0.8 bar, traces every step through OpenTelemetry, exposes itself over A2A, and deploys to Vertex AI Agent Engine. It is the one project in the five-set where the Parallel/Loop primitives genuinely pay for themselves — the other four reserve them — because multi-lens deliberation truly needs Parallel, and the debate and prosecution truly need Loop.
Scaffoldagent-starter-pack
Generated from Google's agent-starter-pack, so the repo is the canonical shape: agent.py exporting root_agent, prompts/ (one per mantri + chair + prosecutor), tools/ (example MCP grounding + A2A clients + the executor's commit/dispatch/file actions), sub_agents/ (five mantris, chair-synthesizer, prosecutor), eval/, tests/, deployment/deploy.py. Built on uv + hatchling.
OrchestrationAgent Development Kit
A coordinator plus a single earned convergence pipeline: ParallelAgent fans out the five mantris simultaneously (the textbook anti-groupthink fan-out); a LoopAgent runs the debate with a deterministic numeric exit; the chair runs as an AgentTool; a second LoopAgent wraps the verdict ↔ Prosecutor gate. Reserved correctly: not a flat agent, not SequentialAgent alone — the deliberation genuinely needs Parallel, the debate and prosecution genuinely need Loop.
Routine intelligenceGemini API
Gemini powers the chair-orchestration and all five advisor mantris — the chair on Gemini 2.5 Pro, the mantris on Gemini 2.5 Flash. Gemini for the many; the volume of deliberation where speed and cost matter.
Decision intelligenceClaude via Vertex AI
The Chair-synthesizer and the Prosecutor run on Claude (Opus-class), deployed through Vertex AI Model Garden — satisfying the challenge's "third-party LLM exclusively through Vertex AI" requirement, and keeping the two highest-stakes reasoning steps (synthesis-under-conflict and beating a steelmanned argument) on a separate, stronger model than the one that produced the positions.
Groundingexample MCP · A2A · Memory Bank
Every advisor argues from the org's own data via tools, never model memory. The opt-in example MCP surface (off until its transport is confirmed; the seed bundle grounds every position meanwhile): admin_stats() → users, analyses, creations, revenue; admin_analytics(type, days) → growth curve, activity/funnel, the churn signal Risk leans on. Cross-org grounding comes over A2A: manas (brand/strategy) and kural (funnel). Long-term grounding — past decisions, standing principles, the founder's revealed preferences — is the roadmap seat for Vertex AI Memory Bank; today continuity comes from manas's versioned, cited Context Pack.
Evaluationchecked-in eval/ @ 0.8
A real eval/ with an evalset.json of past founder decisions (the question, the data available then, the decision made, its known outcome). Two rubric LLM-as-judge tracks at the @0.8 bar: verdict quality (did arivu reach a decision a reasonable board would defend, surface the dissent a human would raise, cite the right grounded numbers) and prosecution soundness (the explicit defensibility ≥ 0.80 gate, a hard pass/fail). A trajectory check confirms grounding tools actually fired and that loop termination was a numeric threshold/escalate, never "advisors agreed." Run in CI via AgentEvaluator before deploy.
Teststests/ — deterministic pieces
A tests/ directory pins the deterministic machinery: the threshold math, the max-iter rollback, and executor atomicity — the parts of the chamber that must behave the same way every time, regardless of what any model says.
ObservabilityOpenTelemetry · BigQuery
OTel tracing spans the whole pipeline — every frame tool call, each parallel advisor, every debate round, the synthesis, each prosecution iteration, the HITL wait, and the executor's actions — so a decision is fully replayable. The BigQueryAgentAnalyticsPlugin streams structured events to BigQuery: decision throughput, rounds-to-converge, confidence distribution, how often the Prosecutor forces a re-verdict (the health metric of the chamber), HITL approve/reject rate, and grounding-tool fire rate. The published resolutions plus the BigQuery log form a durable, auditable record — the company's mind keeping its own minutes.
InteropA2A · to_a2a()
Exposed via to_a2a() with an Agent Card, arivu is both a consumer and a commander in the constellation. It consumes grounding from manas (brand/strategy) and kural (funnel); it commands by dispatching approved verdicts as A2A actions to kural (launch campaign) and kalai (render asset). arivu is the set's only cross-lane node: the single-lane judges inside the executors decide tactics within one lane; arivu decides strategy and priority across all lanes. Opposite data-vector from manas — manas is read-side, arivu is write-side. One crisp seam, no overlap.
Safe actionone gate · no price column
Irreversible action is safe by construction. The pipeline halts at exactly one HITL gate — the named safety anchor; never two gates, never zero. State is resumable (the run registry; Memory Bank is roadmap), so a delayed approval resumes where it paused. And the irreversibility bar is cleared without ever writing a billing/price column on camera — a deliberate retarget given the founder's billing-safety history (RLS column locks, spend_tokens hardening): the committed change is a feature-flag flip or campaign-budget reallocation, not a price/revenue write. The downstream A2A dispatch causes the real, irreversible spend.
DeployVertex AI Agent Engine · Cloud Run
The chamber runs on Vertex AI Agent Engine as the managed primary runtime — the convergence pipeline, the HITL pause/resume, and Memory Bank all run natively there, with tracing wired to the OTel spans above. Cloud Run is the secondary target for the A2A endpoint front door if needed. Both ship from deployment/deploy.py — all native Google Cloud. GKE appears in zero ADK samples and is never used.
The north stars — the discipline that defines the product
01
DECIDE, don't recommend
arivu always terminates in either a committed, dispatched, on-the-record decision — or an explicit "no safe decision — re-frame." A recommendation that leaves the human to act is a failure mode, not the product.
02
Grounded or silent
No advisor speaks from model memory. Every position cites a number pulled live from the org's own data via a tool, or it is not admitted to the chamber.
03
Survive the prosecution
A verdict ships only after it has beaten its own steelmanned null case at defensibility ≥ 0.80. The chamber prosecutes itself before it ever asks the human to trust it.
04
One gate, one tap
Exactly one human approval stands between deliberation and irreversible action — the named safety anchor — and it is unmissable. Never two gates, never zero.
05
Preserve the dissent
The minority position is recorded in the board resolution, never erased. The company can always see what it decided against, and why.
06
Deterministic termination, always
Every loop ends on a numeric threshold or a max-iter rollback — never on "the advisors agreed." No runaway deliberation, ever.
05
Business case · impact
The board you couldn't hire.
The impact isn't a vanity metric — it's the quality of the company's decisions and the speed at which they become reality. arivu turns the strategic-deliberation work that no four-person company can staff into a standing capability: a grounded board, a chair, and an adversary, convened on demand and bound to one approval.
Take Sundara Coffee Co. The founder asks the chamber, "Should we raise Pro to $39?" The Economist pulls live revenue; Growth pulls the funnel; Brand pulls the canon; Risk pulls churn cohorts and catches a cliff past $36; Ops confirms the billing flag is safe to flip. The Claude chair strikes "Raise to $34, grandfather existing users, 30-day notice," the Prosecutor fails to defeat it at defensibility 0.84, and the founder taps once to approve. The flag flips, kalai renders the banner, and kural — the company's only mouth — launches the announcement and carries the banner to the world behind a second, downstream publish sign-off at the mouth. A signed resolution lands at a real URL, dissent and all. A 1am gut call became a defensible, executed, on-the-record decision.
Solo founder
Stops deciding alone at 1am
The loaded calls that used to ride on instinct now go to a grounded chamber that argues from real numbers, prosecutes itself, and hands back one defensible decision — with the dissent on the record.
Small team
Skips the fractional-CFO hire
A CFO, a growth lead, a brand head, and a devil's advocate as software — the board you couldn't justify hiring yet, convened on demand for the cost of model tokens, and bound to your single approval.
The company
Decisions become reality, on the record
Every call is grounded, prosecuted, approved, executed across every team, and filed as a signed resolution. The company can replay exactly what it decided, why, and what it decided against — its own minutes, kept automatically.
≥ 0.80
defensibility a verdict must beat to ship
1 tap
between deliberation and irreversible action
5 + 2
grounded lenses · Claude chair & prosecutor
100%
of decisions filed as a signed resolution
Raises decision quality. Five disjoint grounded lenses catch the blind spot — the churn cliff — that one model answering from memory ships straight past.
Makes decisions defensible. Nothing reaches the human until it has survived its own steelmanned prosecution at defensibility ≥ 0.80 — the company never trusts a verdict the chamber couldn't.
Closes the execution gap. A verdict doesn't end as advice; it commits a live change and dispatches real orders across kural and kalai — the decision becomes action in lockstep.
Keeps the minutes. Every call is published as a signed board resolution with preserved dissent and a content hash — a durable, auditable record of every decision the company has made.