# Agent Shipyard

> **About the citations:** every outside document this report cites was copied into an archive folder, `canon/`, on 2026-07-05. The citations show what the report read when it was written and haven't been updated, so to open one now, look for its copy under `canon/`.

:::animation HERO
**HERO: The shipyard beside the marketplace**
- **What it shows:** a vast harbor at dusk; on the left a working shipyard where agents (rendered as vessels) are fitted out with instrumentation, run through evaluation dry-docks, fuel-gauged for cost, and stamped with governance seals; on the right, directly beside it, a bustling marketplace where those same vessels are traded, each carrying a visible trust-manifest (reliability score, cost-per-task, a verified-creator flag) that was earned in the shipyard; goods flow from the works into the market in one continuous motion.
- **Narrative role:** sets the scene; this is the share/card thumbnail. It frames Agent Shipyard as the junction where agents are made seaworthy and then traded, with the operational telemetry from the works becoming the trust signals of the market.
- **What it teaches:** the one idea is that operation and commerce belong together, because an agent's observed reliability, cost, and quality are exactly what let a buyer trust it.
- **Intended impact:** the realization that the fragmented agent-ops tools and the trustless agent marketplaces are two halves of one missing whole.
:::

| Field | Value |
|---|---|
| Project | Agent Shipyard |
| Looikos cluster | Infrastructure & Agent Platforms (the infrastructure-and-monetization layer of agents) |
| One-line | The place where agents are deployed, observed, evaluated, budgeted, governed, and monetized, and the marketplace where they are bought and sold, the shipyard beside the great marketplace where the goods come in |
| Status | Concept (depends on the harness and the design and tooling layers it operates on top of) |
| Existing code | None yet; sits on top of Symphony AGI (the harness the agents run on), beside MCP Scientists (the tools the agents use) and Agent Design Pro (where the agents are made) |
| Desk | desk-infra (written by desk-brands-finish) |
| Coverage | VERIFIED-heavy on the seed (Andy's transcript is first-party and current) and on the AgentOps market and the build reality (Perplexity, cited, with real market figures and the OTel/eval/cost/marketplace stack). INFERRED on the three-angle valuation comps (shared agent-infra category) and persona PST depth (tagged inline) |
| Date | 2026-06-21 |

---

## Nine-rung frame (this research task)

- **Purpose (the rails):** give the ecosystem the depth to build and run Agent Shipyard with agents, not headcount. Agent Shipyard is the layer that makes the ecosystem's fleet of agents observable, affordable, and trustworthy, which is the precondition for running dozens of brands without a department watching each one.
- **Mission (rung 1):** convert Andy's recorded Agent Shipyard breakdown into a research-grounded ~10k brand deck, so the agent-infrastructure-and-monetization brand is designed and sold from understanding the operational pain, not from a logging-vendor's-eye view.
- **Objective (rung 2):** a finished deck at `symphony/stack-recon/projects/agent-shipyard.md`, ~10k words, three-angle valuation modeled, 5+ PST personas to world-experience depth, build section grounded in the agent-ops and marketplace build reality, graded CLEAN by desk-qc-final and the lead.
- **Initiative (rung 3):** the symphony-recon Track-P run; one of the ~11 unwritten decks.
- **Project (rung 4):** the desk-brands-finish lane.
- **Task (rung 5):** this one brand deep-dive, run against `_PROJECT_TEMPLATE.md` and PST.
- **Action (rung 6):** A1 ingest the transcript (done). A2 skeleton (done). A3 sequential Perplexity: AgentOps market/alpha, Lexicon of Pain, build reality (done, three queries; valuation comps reused from the same-category runs). A4 PST on five personas. A5 incremental section writing. A6 self-check. A7 hand to the lead.
- **Decision (rung 7):** the evolution stage of the agent-ops capability (basic observability is heading to commodity on the OpenTelemetry standard; the unified deploy-observe-evaluate-budget-monetize control plane plus the trust-signal marketplace is the genesis-stage own-it lane); which personas carry the deck (the blind-in-prod engineer, the cost-bankruptcy-fearing team, the unmonetized builder, the distrustful buyer, plus the governance-bound enterprise); ADOPT/HARVEST on the build (adopt OTel GenAI conventions and the open eval/observability stack, own the unified control plane and the trust-signal marketplace and the cost-governance-as-first-class layer).
- **Data (rung 8):** N/A (this doc is the artifact). It later seeds the metagraph as BrandDeck:AgentShipyard, with edges to Symphony AGI, MCP Scientists, Agent Design Pro, and Swarm Layer.
- **Event (rung 9):** N/A (this doc is the artifact). Deck written, progress posted, grade recorded.

## 1. What it is (the one-paragraph truth)

Agent Shipyard is the infrastructure and monetization layer of agents: the place where agents, once designed and built, are deployed, run, watched, measured, budgeted, governed, and sold. In Andy's words it is where you "deploy and manage the ecosystem of available agents, the marketplace of agents," distinct from where the tools agents use are made (MCP Scientists) and where the agents themselves are designed (Agent Design Pro). Concretely, it fuses two things. The first is an agent operations control plane: deployment management across environments, historical execution logging, analysis of input and output and tool-use patterns, evaluation, prompt backtesting and split testing, cloud-budget and token-usage optimization, and the management of integrations, dependencies, and privacy settings, all in one place. The second is a marketplace: a place where available agents are listed, discovered, compared, bought, and deployed, by enterprises and by everyday people, with the operational telemetry from the first half feeding the trust signals of the second half. The name carries the strategy. A shipyard is where vessels are built, fitted, repaired, and launched, and the great marketplaces of history always sat beside the shipyards because that's where the goods come in. Agent Shipyard is positioned at that junction: the operational works where agents are made seaworthy, beside the market where they are traded, with the two reinforcing each other because an agent's observed reliability, cost, and quality become the signals that let a buyer trust it. For the people it serves, it answers a set of miseries that the fragmented tooling market leaves unsolved: engineers flying blind in production, teams terrified of a runaway-cost bill, builders who made something useful and earn nothing from it, and buyers who can't tell a working agent from a scam.

:::animation 1
**ANIMATION 1: Two halves fused into one**
- **What it shows:** two separate panels, an "operations control plane" (deploy, observe, evaluate, budget, govern) and a "marketplace" (list, discover, compare, buy), slide together and lock; a stream of telemetry from the ops panel flows directly into the marketplace as trust signals (a reliability score, a cost-per-task, a compliance badge) appearing on each listing.
- **Narrative role:** opens §1 by making the brand's defining fusion (ops plus marketplace, telemetry as trust) immediately legible.
- **What it teaches:** that the two halves only work fused, because the operational data is what makes the market trustworthy.
- **Intended impact:** the viewer sees why the brand is one connected whole, not two adjacent products.
:::

:::animation 1b
**ANIMATION 1b: four miseries the fragmented market leaves**
- **What it shows:** four figures stand stranded across a broken landscape, an ENGINEER FLYING BLIND squinting at a black box, a TEAM watching a cost meter spike toward a runaway bill, a BUILDER holding a working agent with an empty till beside it, and a BUYER unable to tell a real agent from a scam, each stranded by a different gap in the same market
- **Narrative role:** anchors the closing claim of §1, the set of miseries the fragmented tooling market leaves unsolved
- **What it teaches:** the brand exists to answer four concrete pains at once, not to add one more logging vendor
- **Intended impact:** the reader holds the four buyers the brand is built to rescue before meeting them as personas
:::

## 2. Andy's seed, expanded

**Andy's words (verbatim from the recording):** "Agent Shipyard is next. So pretty much what we discussed as MCP scientists being for tooling or agentic tooling, so MCPS tools for agents. Agent Shipyard is where we actually deploy and manage the ecosystem of available agents, the marketplace of agents. And so you can look at it as. There's another brand we'll talk about soon called Agent Design Pro. That's where agents are actually made. But here with Agent Shipyard, Agent Shipyard is where you would do things such as managing deployment, looking at historical executions. Obviously we save everything. So you can look at patterns of input, output, tool use. You can see evaluation and run prompt back tests and split tests and manage different deployments across different environments and manage your cloud budget and optimize based on token usage and apply different types of integrations and dependencies and privacy settings. There's a ton of stuff that goes into Agent Shipyards or consider it like the infrastructure layer of agents, the infrastructure and monetization layer of agents. This is also where your everyday Joe Blow, that's where we think like a shipyard. Whereas imagine the massive marketplaces that are all always close to shipyards. There's always going to be a demand for that because this is where the goods are coming in."

:::animation 1a
**ANIMATION 1a: the name carries the strategy**
- **What it shows:** a working shipyard sits on the water and, right against its gates, a crowded marketplace grows; goods roll off the slipways straight onto the market stalls, a label reading THIS IS WHERE THE GOODS COME IN pinned to the gate between them, the two never separated
- **Narrative role:** anchors the §1 claim that the shipyard-beside-the-marketplace name is the strategy, not decoration
- **What it teaches:** great markets always sat beside the works because the works are where the tradeable goods arrive
- **Intended impact:** the reader reads the brand name as a positioning argument rather than a metaphor
:::

**Reading between the lines:** Andy's seed compresses three claims, and each one is load-bearing.

First, the boundary against its two siblings is stated explicitly, and it's the discipline that makes the family of brands coherent rather than one undifferentiated agent-platform. MCP Scientists makes the tools agents call. Agent Design Pro is where agents are made. Agent Shipyard is where made agents are deployed and operated and sold. Keeping these three layers distinct is itself a thesis, because the loudest operational complaint in the 2026 agent market is the tangle of half-overlapping tools where deployment, observability, evaluation, cost, and distribution are split across many vendors and nothing connects (VERIFIED, Query 1: the observability vendors do visibility not distribution, the runtime hosts do hosting not marketplace dynamics, the marketplaces do discovery not telemetry, and the full lifecycle in one place is exactly the gap). Andy's instinct to draw the line cleanly is the instinct to own the connected whole that the market leaves fragmented.

:::animation 2a
**ANIMATION 2a: three layers, one clean line each**
- **What it shows:** three labeled bands stack without overlapping, MCP SCIENTISTS makes the tools agents call, AGENT DESIGN PRO makes the agents, AGENT SHIPYARD deploys and operates and sells them, and where the market's incumbents blur into a tangled knot beside them, these three hold a clean seam
- **Narrative role:** anchors the first claim of §2, the explicit boundary against the two sibling brands
- **What it teaches:** drawing the three layers apart is itself the thesis that lets the brand family stay coherent
- **Intended impact:** the reader sees the discipline of the clean boundary against the market's half-overlapping tangle
:::

Second, the enumerated capability list is, read carefully, a precise description of the modern AgentOps control plane, and the order Andy lists them in tracks the actual agent lifecycle. "Managing deployment" and "different deployments across different environments" is the deployment and runtime layer. "Looking at historical executions, we save everything, patterns of input, output, tool use" is the observability and tracing layer, which in 2026 is built OpenTelemetry-first on the GenAI semantic conventions, with sessions and runs and model spans and tool spans and token-usage metrics (VERIFIED, Query 3). "Evaluation and run prompt back tests and split tests" is the evaluation and online-experimentation layer (golden datasets, LLM-as-judge, regression gates in CI, shadow deployments, and live split testing across prompt and model and agent variants). "Manage your cloud budget and optimize based on token usage" is cost governance (per-agent and per-tenant cost attribution, budget guardrails, circuit breakers, loop detection, model routing for cost). "Integrations and dependencies and privacy settings" is the governance layer. Every item Andy names maps onto a real, documented piece of the 2026 stack, which makes the brand a productization of a real and forming category, not aspirational hand-waving.

:::animation 2
**ANIMATION 2: The agent lifecycle as a control plane**
- **What it shows:** an agent's lifecycle laid out as a horizontal track, each station lighting up in sequence (deploy across environments, observe via OpenTelemetry spans, evaluate with golden-dataset gates, budget with circuit breakers, govern with policy spans), each station revealing the real 2026 tooling underneath it; Andy's enumerated capability list maps onto the stations one to one.
- **Narrative role:** illustrates §2's second claim, that Andy's list is a precise description of the modern AgentOps control plane.
- **What it teaches:** that the brand productizes a real, documented, forming category rather than inventing one.
- **Intended impact:** the viewer sees the brand is grounded in a live technical reality.
:::

Third, "the infrastructure and monetization layer of agents" plus "your everyday Joe Blow, we think like a shipyard, the marketplaces always close to shipyards, this is where the goods are coming in" is the move that separates Agent Shipyard from every pure-ops tool: it ties the operations to the market. The observability and evaluation data that the ops layer produces is exactly the trust signal a marketplace needs, and no incumbent does this. The observability tools help teams understand behavior but don't turn usage, quality, and reliability data into marketplace ranking, revenue share, or creator payouts; the marketplaces (OpenAI's GPT Store, Salesforce AgentExchange, the cloud agent catalogs) handle discovery and distribution but carry no deep operational telemetry or trust signals (VERIFIED, Query 1 and Query 4). Andy's "shipyard beside the marketplace" is the architectural claim that the operational works and the market belong together, because the operational data is what makes the market trustworthy, and a trustworthy agent market is exactly what the distrustful everyday buyer is waiting for. The "everyday Joe Blow" detail points at a second piece of open ground: the existing marketplaces are ecosystem or enterprise distribution channels, not a broad consumer-style market where ordinary people buy a working agent for a specific job.

:::animation 3
**ANIMATION 3: Operations feeding the market**
- **What it shows:** the ops control plane on the left continuously producing observability and evaluation data; that data streams rightward and crystallizes into marketplace trust signals (ranking, revenue share, creator payouts) on the listings; the incumbents are shown as two disconnected islands (observability tools that do not monetize, marketplaces that carry no telemetry) with a dark gap between them that Agent Shipyard bridges.
- **Narrative role:** illustrates §2's third claim, the operations-feeds-the-marketplace move that no incumbent makes.
- **What it teaches:** that the operational data is the trust signal the marketplace needs, and bridging the gap is the alpha.
- **Intended impact:** the viewer sees the structural opening between the two incumbent islands.
:::

## 3. The three-angle valuation (the core of a self-standing brand)

Agent Shipyard stands on all three angles with a distinctive shape: its finance angle is the strongest in the family because a marketplace with a take-rate plus usage-metered ops revenue is a payments-and-throughput business that the capital markets value richly; its software angle is a genuine control plane plus a two-sided marketplace; and its service angle is the managed-operation layer for teams that can't run their agent fleet themselves.

### 3a. Finance (credit and capital access)

The category read here comes with real market figures. The AI agents market is put at roughly $7.6B in 2025 and $10.9B in 2026, growing at a 44-46% CAGR toward $182.9B by 2033 (VERIFIED, Query 1, Grand View Research and corroborating sources). Agent Shipyard sits at the operational-and-commerce center of that market, and the M&A context from the Symphony AGI deck on this site applies: roughly 90% year-over-year growth in AI M&A deals and a $500M-$5B strategic-valuation band for integrable agentic-AI assets (VERIFIED, valuation query, referenced from `projects/symphony-agi.md` §3a). Two sets of comparables matter most here. On the ops side, they're the bands of multiples on annual recurring revenue (ARR) for dev-tooling and AI-infra companies (top-tier AI-infra 12-20x, dev-tools 8-12x, mid-tier 4-8x) and the funded observability vendors (LangSmith, Langfuse, Arize, Braintrust). On the marketplace side, the more interesting comparable is the marketplace-business valuation logic: a two-sided market with a take-rate on gross transaction value is valued on marketplace multiples (roughly 2-5x GMV or 10-20x marketplace revenue, the band for businesses like the Stripe or Shopify app stores) rather than pure SaaS multiples, especially when there are real two-sided network effects where more agents attract more buyers and vice versa (VERIFIED, valuation query). Agent Shipyard is the one brand in the family whose finance angle benefits from being read partly as a marketplace, which is a higher-multiple lens than pure infrastructure.

How that converts to credit and capital access has a marketplace-specific nuance. The ops-subscription and managed-operation retainer revenue is contracted and recurring (the good collateral, underwritten like classic SaaS: revenue-based financing at 20-40% of ARR to a 1.2-1.5x cap, ARR-backed debt at 0.3-0.8x ARR, 8-15% plus warrants). The marketplace revenue is the transactional line, and lenders treat it carefully: if a large share of the gross transaction value is recurring automation (agents running ongoing workflows for enterprise buyers), lenders may underwrite a portion of it as quasi-ARR; if it is one-off task purchases, they discount it heavily as transactional (VERIFIED, valuation query). The strategic implication is to design the marketplace toward recurring agent subscriptions rather than one-off purchases, because that converts the GMV from heavily-discounted transactional revenue into quasi-recurring revenue with real credit capacity, and it also builds the two-sided retention that the marketplace multiple rewards. The accumulated proprietary state that an acquirer pays the strategic premium for is unusually rich here: the standardized telemetry across every agent that runs on the platform, the trust-and-ranking models trained on that telemetry, the creator and buyer relationships, and the operational track record, none of which a competitor can clone by listing the same agents.

:::animation 3a1
**ANIMATION 3a1: design the market toward recurring, and it borrows**
- **What it shows:** a marketplace of one-off task purchases, each sale a coin that drops and vanishes, gets redesigned toward recurring agent subscriptions, and the same sales become a steady contracted stream a lender marks as quasi-recurring; a credit line then opens against that stream while the one-off pile stays discounted and dark
- **Narrative role:** anchors the §3a credit-conversion claim, the recurring-versus-transactional GMV nuance
- **What it teaches:** steering the marketplace toward recurring subscriptions converts heavily-discounted transactional revenue into borrowable quasi-ARR
- **Intended impact:** the reader sees a product-design choice reprice directly into credit capacity
:::

The market-maker read runs on three levels. On fundamentals, it's a fast-growing, well-quantified market with Agent Shipyard at its operational-and-commerce center and a defensible position that pairs a two-sided market with telemetry. On technicals, the supply of unified control planes is thin (everyone does one layer), and a credible neutral, model-agnostic, self-hostable one is rarer still, against broad demand, which is a favorable order book. On sentiment, agent operations and governance is the consensus 2026 priority as agents move from pilot to production, with the named risk that the hyperscalers (Bedrock AgentCore, Vertex Agent Engine, Azure AI Foundry) fold ops into their clouds; the hedge is the neutral, model-agnostic, self-hostable, cross-cloud positioning plus the consumer-marketplace ground the hyperscalers don't serve.

:::animation 4
**ANIMATION 4: The marketplace multiple versus the SaaS multiple**
- **What it shows:** two valuation dials side by side; a "pure SaaS infrastructure" dial reading a solid multiple, and a "two-sided marketplace with take-rate and network effects" dial reading higher (the 2-5x GMV / 10-20x marketplace-revenue register); more agents flowing in spin the marketplace dial up as buyers follow, illustrating the two-sided network effect that lifts Agent Shipyard's finance angle.
- **Narrative role:** grounds §3a's distinctive marketplace-valuation logic.
- **What it teaches:** that Agent Shipyard's finance angle is read partly as a marketplace, a richer lens than pure infrastructure.
- **Intended impact:** the viewer sees why this brand's finance angle is the strongest in the family.
:::

:::animation 3a2
**ANIMATION 3a2: the moat an acquirer cannot clone**
- **What it shows:** a rival tries to copy Agent Shipyard by listing the identical agents, and the listings appear, but the proprietary state stays behind glass on the original, the STANDARDIZED TELEMETRY across every run, the TRUST-AND-RANKING MODELS trained on it, the CREATOR AND BUYER RELATIONSHIPS, and the OPERATIONAL TRACK RECORD, none of it transferring with the copied catalog
- **Narrative role:** anchors the §3a claim that the strategic premium rests on accumulated proprietary state
- **What it teaches:** the telemetry, the trained trust models, and the two-sided relationships are what a competitor cannot reproduce by listing the same agents
- **Intended impact:** the reader sees where the durable value sits, in the accumulated operational state rather than the catalog
:::

### 3b. Software (the interface stack)

Agent Shipyard's software is a four-layer control plane plus a marketplace, and each layer maps onto a real, documented 2026 architecture (VERIFIED, Query 3 and Query 4).

The telemetry layer is OpenTelemetry-first on the GenAI semantic conventions: a hierarchical trace model of session spans, run spans, model-call spans (carrying gen_ai.request.model and gen_ai.usage token metrics), tool spans (including the MCP spans added in OTel 1.39), retrieval spans, and guardrail spans, collected through an OTel Collector with processors for PII scrubbing, cost enrichment, and per-tenant tagging, and stored in a high-volume backend (ClickHouse-class, the kind Langfuse uses). The evaluation-and-experimentation layer runs golden-dataset regression suites in CI as pre-deploy gates, LLM-as-judge and rule-based evaluators online on sampled live traffic, trajectory monitors that catch the subtle "still kind of works but the workflow changed" regressions, shadow deployments, and live split testing that routes traffic across prompt and model and agent variants with the variant IDs stamped on every trace so metrics slice by variant. The cost-governance layer enriches every model span with a cost_usd from pricing tables, aggregates by tenant and agent and feature, enforces budget guardrails and circuit breakers (max spend per user per day, loop detection via trajectory monitors, max tool calls per session), and routes models for cost (cheap-versus-premium selection driven by the observed quality-versus-cost curve). The governance layer handles integrations, dependencies, privacy settings, and the audit trail.

The marketplace layer is the differentiator, and it's built on the telemetry: a catalog of agents as versioned entities with creator identity and required permission scopes; a sandboxed runtime that enforces instrumentation so every listed agent produces standardized telemetry; an eval-and-ranking engine that turns that telemetry into public trust signals (quality scores, reliability SLAs, cost-per-task, usage and retention, compliance badges) and into the ranking that demotes misbehaving or degraded agents quickly; and a metering-and-payouts layer in the Stripe-Connect register (the marketplace is the platform merchant, creators are connected accounts receiving usage-based revenue share, the platform holds funds and handles refunds and retains its fee) (VERIFIED, Query 4). The product surfaces and monetization decompose cleanly: the control plane is the SaaS (subscription by agents-under-management, environments, and enterprise features); the ops capabilities are usage-metered where they scale with volume; the marketplace is the take-rate business (a percentage of gross transaction value, in the 10-30% range); and the trust-signal certification can itself be a premium tier. The architectural rule that ties it to the ecosystem is that it's model-agnostic and self-hostable, the neutral position the cloud-locked incumbents can't occupy, and that it runs the ecosystem's agents (the fleet running on the Symphony AGI harness) as its first and most demanding tenant.

:::animation 5
**ANIMATION 5: The four-layer control plane plus the marketplace**
- **What it shows:** a vertical stack building up layer by layer, telemetry (OTel spans), evaluation-and-experimentation (golden-dataset gates, split tests), cost-governance (circuit breakers, model routing), governance (audit trail), and then the marketplace layer crowning it (catalog, sandboxed runtime, eval-and-ranking, Stripe-Connect payouts), the marketplace visibly drawing its trust signals down from the telemetry at the base.
- **Narrative role:** grounds §3b's software architecture in the real documented stack.
- **What it teaches:** that the marketplace is built on the telemetry, which is what makes it different from a listings store.
- **Intended impact:** the viewer sees the full architecture and why the marketplace rests on the ops layers.
:::

:::animation 3b1
**ANIMATION 3b1: one platform, four ways it earns**
- **What it shows:** the platform core splits its revenue into four labeled channels, CONTROL-PLANE SaaS billed by agents-under-management, OPS CAPABILITIES metered by volume, the MARKETPLACE TAKE-RATE at a slice of gross transaction value, and a TRUST-SIGNAL CERTIFICATION premium tier, each channel a different-shaped meter running off the same underlying platform
- **Narrative role:** anchors the §3b monetization decomposition, how the product surfaces map to revenue
- **What it teaches:** the same platform earns four distinct ways, a subscription, a meter, a take-rate, and a certification tier
- **Intended impact:** the reader sees a diversified revenue base rather than a single pricing line
:::

### 3c. Service (premium-at-accessible boutique delivery)

The service angle is the managed-operation layer: running an agent fleet on behalf of teams that have agents in production and no ability to operate them well. The target operator is the sub-25-employee company or the team that shipped agents and is now flying blind, bleeding cost, and afraid of the next deploy. The service delivers premium quality at accessible prices through the pre-built control plane: because Agent Shipyard already has the telemetry, the evals, the cost guardrails, and the governance built, it can drop a client's agents onto an observable, evaluated, budgeted, governed footing at a fraction of what building that operational stack would cost, and faster.

The retainer economics follow the ecosystem standard: $1-2k accessible at entry, $2-12k+ for the real engagements, structured as a managed-operations retainer (we deploy, observe, evaluate, budget, and govern your agents) plus overage for scale. A target of 100-250 customers puts the service angle's floor around $1M/month, and it scales above that (VERIFIED, `THE_FLOOR.md`). The trust differentiator answers the two deepest operational fears in the Lexicon of Pain (the phrases people use when they describe the problem themselves): the fear of the runaway-cost bill that "could quietly bankrupt us," answered by real budget guardrails and circuit breakers and per-tenant attribution; and the fear of silent regressions and blind production, answered by the OTel-first tracing and the continuous eval monitors. The ongoing fleet operation and the long-tail support get handed to the ecosystem's affiliate network, which runs on a shared-floor model: senior technical people in emerging markets work through the Looikos tools with a franchise-style path to ownership, and live transcripts and agent-native systems let that floor run globally (VERIFIED, `THE_FLOOR.md`). The vertical doesn't matter; any team running agents in production that needs them observed, evaluated, budgeted, and governed qualifies, and the service angle is also the on-ramp that feeds the marketplace, because a client whose agents are well-operated is a client whose best agents can be listed with real trust signals.

:::animation 6
**ANIMATION 6: The service on-ramp into the marketplace**
- **What it shows:** a client's chaotic, un-operated agents enter the managed-operations service on the left, get observed, evaluated, budgeted, and governed (the chaos resolving into clean instrumented vessels), and then the best of them sail directly into the marketplace carrying the trust signals they earned during the managed operation.
- **Narrative role:** grounds §3c's service angle and its connection to the marketplace.
- **What it teaches:** that the managed-operation service is also the pipeline that feeds well-operated, trust-signal-bearing agents into the market.
- **Intended impact:** the viewer sees the service and the marketplace reinforcing each other.
:::

:::animation 3c1
**ANIMATION 3c1: the two operational fears, answered**
- **What it shows:** a team carries two dreads into the managed service, a RUNAWAY-COST BILL looming overhead and a SILENT REGRESSION creeping through a blind production system; the service meets each one directly, budget guardrails and circuit breakers and per-tenant attribution capping the bill, OTel-first tracing and continuous eval monitors lighting up the blind spot, and both dreads shrink
- **Narrative role:** anchors the §3c trust differentiator, the two deepest operational fears the service answers
- **What it teaches:** the managed layer earns trust by answering the bankruptcy-bill fear and the blind-regression fear with built machinery
- **Intended impact:** the reader sees the service selling relief from two specific, named fears rather than generic operations
:::

## 4. The personas (5+, modeled to world-experience depth)

The language here is pulled from the actual Lexicon of Pain mined in the voice-of-customer research (the second research query). Agent Shipyard serves both sides of the market, so the personas span the producers (engineers and builders who deploy and want to monetize) and the consumers (teams and everyday people who buy), and the shame textures differ: the producer's shame is "I'm the LLM guy and I'm basically guessing," the consumer's is "I don't even know what questions to ask."

:::animation p0
**ANIMATION p0: two sides of one market, two shames**
- **What it shows:** a market splits down the middle, PRODUCERS on one side (engineers and builders who deploy and want to earn) and CONSUMERS on the other (teams and everyday people who buy), the producer wearing a thought reading I am the LLM guy and I am basically guessing, the consumer wearing one reading I do not even know what questions to ask, the same trust gap running between them
- **Narrative role:** frames the whole persona section, the producer-and-consumer split and the two shame textures
- **What it teaches:** the five personas divide into producers and consumers, and each side carries a distinct shame the brand must answer
- **Intended impact:** the reader reads the personas as two halves of one market rather than five unrelated buyers
:::

### Persona 1: The engineer flying blind in production

I'm the engineer who shipped a "smart" agent and now lives in dread of it. "My agent is basically a black box once it hits production. I have no idea what it's doing step-by-step, just occasional 500s and some vague logs." Multi-step runs "disappear into the void," and when something fails "all I get is tool call failed with zero context." I "can't tell which step of the agent chain is actually breaking," it's "like debugging with a blindfold on." Right now my observability is "tailing logs and grep, that's it." And the regressions are silent: "we tweak one prompt and two weeks later support tickets spike because the agent quietly got dumber in one edge case," with "no alerts, no tests, nothing." I'm "terrified of touching the prompts now," everything is "held together with vibes and duct tape," and every release "is a gamble, we don't ship features, we ship experiments on our users."

:::animation p1v
**ANIMATION p1v: debugging with a blindfold on**
- **What it shows:** an engineer stands over a production agent that is a sealed black box, only occasional 500s and vague logs leaking out; a multi-step run disappears into a void, TOOL CALL FAILED prints with zero context, and the engineer's only instrument is a terminal scrolling tail -f logs, everything visibly held together with vibes and duct tape
- **Narrative role:** carries persona 1's first-person voice, the raw texture of flying blind in production
- **What it teaches:** the pain is operating an opaque black box with no instrument better than grep, so every release is a gamble
- **Intended impact:** the reader feels the specific dread of shipping experiments on users without being able to see the run
:::

It hits me where my professional identity lives. I'm "supposed to be the LLM guy, but I'm basically guessing," and the shame is "fake robustness," presenting this thing as AI automation when I know it's duct tape. What I fear is hidden failure ("I don't actually know when it's wrong") and getting blamed ("when it screws up in front of a customer, it's on me and I can't even explain what happened"). I ended up here because shipping the agent was the goal, and operating it was an afterthought the tooling didn't support. Getting out takes real traces (every thought, tool call, and model response as a nested span with timing and tokens), evaluation gates that catch a regression before it ships, and continuous monitors that alert me when quality drops. Most engineers in my seat fail because they keep adding more printf debugging instead of adopting an observability stack. Staying stuck costs me the hidden failures and the blame, and getting out costs me admitting that "tail -f logs" isn't observability and standing the agent on a real control plane.

What Agent Shipyard offers me is the thing I've been faking: OpenTelemetry-first tracing so I can finally see the whole run, evaluation gates so a prompt change is red or green before it ships, and monitors that tell me when the agent quietly got dumber, so I stop shipping experiments on my users.

:::animation 7
**ANIMATION 7: Blindfold off**
- **What it shows:** an engineer debugging by scrolling a flat wall of grey log text (blindfolded), then the blindfold lifts and the same run renders as a nested span tree (each thought, tool call, and model response with timing and tokens), the failing step glowing red and immediately identifiable; a regression alert fires the moment a prompt change degrades a metric.
- **Narrative role:** dramatizes persona 1's transformation from blind log-tailing to real observability.
- **What it teaches:** that real tracing turns an opaque black box into a legible, debuggable run.
- **Intended impact:** the viewer feels the blindfold come off.
:::

### Persona 2: The team being bankrupted by runaway cost

I'm the infrastructure person on a team that "woke up to a 5-figure OpenAI bill because an agent looped overnight, no alert, no guardrails, just surprise." Token spend "is a complete black box," and when finance asks "why did this month triple" I have "no answer except the agent talked too much." I "can't answer the most basic question: which customer or which agent is actually costing us money," because "all I get is one giant LLM usage number." There's "no circuit breaker for agents," I "want to say this user can't spend more than five dollars a day and the tooling just doesn't exist," and "our only real guardrail is me hovering over the usage dashboard and praying nothing spikes." "I'm more scared of the usage bill than of the bugs at this point," because "every deploy feels like it could quietly bankrupt us if we screw up a loop condition."

It hits my standing and my sense of responsibility. The shame is being "the infra person" who "can't tell finance where the money is going," and feeling like I'm "running a production system with no rate limits and hoping no one notices." The fear is sharper: a dumb prompt mistake that "could literally cost real cash we don't have," and the knowledge that "if this happens again, they're going to shut down the whole AI initiative." I'm here because an agent's costs are invisible by default and the guardrails aren't built in. What would get me out is per-agent and per-tenant cost attribution so I can answer finance, budget guardrails and circuit breakers so a loop can't bankrupt us, loop detection that cuts off a runaway agent, and model routing so the cheap model handles what it can. Most teams fail because they bolt on a manual hard limit and hope. If I stay stuck, I pay with the surprise bill and the killed initiative. Getting out means treating cost governance as a first-class system instead of a dashboard I stare at.

What Agent Shipyard offers me is cost governance as a real product: every model call cost-attributed by tenant and agent, hard budget guardrails and circuit breakers that stop the loop before it stops us, and routing that keeps the bill rational, so I can finally tell finance exactly where the money goes.

:::animation 8
**ANIMATION 8: The circuit breaker that stops the loop**
- **What it shows:** an agent spiraling into a runaway loop overnight, the token-cost meter climbing toward a five-figure surprise bill, then a circuit breaker trips (a trajectory monitor detecting the repeated tool call) and halts the loop; the cost is then attributed cleanly by tenant and agent in a panel a finance lead can read.
- **Narrative role:** dramatizes persona 2's transformation from runaway-cost terror to governed spend.
- **What it teaches:** that hard guardrails and per-tenant attribution turn an unbounded risk into a controlled, explainable cost.
- **Intended impact:** the viewer feels the bankruptcy fear defuse.
:::

### Persona 3: The builder who made a great agent and earns nothing

I'm the indie developer who "built a GPT that users actually love, hundreds of messages, great feedback, revenue: zero." The store "was sold as the App Store for AI and it's basically a random list of spammy bots and SEO bait." "Discovery is nonexistent, unless they bless you in featured, you're invisible," and I'm "competing with 10,000 copy-paste SEO GPTs and crypto shills." It's "lottery economics, a tiny handful of creators make real money, everyone else is just adding free content to their platform." Worse, "you don't get any of the user relationship, no email, no way to talk to your customers," and "I'm locked into their store with zero portability." It's "like building a SaaS on top of a slot machine, they can change the odds whenever they want." I built something that "actually saves people hours," but "there's no way to charge per seat, per team, or integrate with their workflows," so "the ceiling is like a hundred dollars a month, that's not a business, that's a tip jar."

:::animation p3v
**ANIMATION p3v: a SaaS built on a slot machine**
- **What it shows:** a builder's genuinely-loved agent sits inside a closed store drawn as a slot machine, the odds dial owned by the platform and changeable at will; DISCOVERY IS NONEXISTENT flickers across a wall of copy-paste spam, the user relationship is sealed away with no email and no portability, and the revenue counter is frozen at a hundred-dollar tip-jar ceiling
- **Narrative role:** carries persona 3's first-person voice, the raw texture of building on an extractive closed store
- **What it teaches:** the pain is lottery economics with no portability, where good work earns a tip jar because the platform owns the odds
- **Intended impact:** the reader feels why the builder needs a market that pays on merit rather than a blessing
:::

It hits me as someone trying to make a living from real work. The shame is "feeling naïve," that "I bought the App Store for AI dream and ended up giving them free labor," and not having "a real startup," just being "an AI creator on someone else's platform." The fear is "building on shifting sand" where "any day they can change the rules and kill my income." I landed here because the only distribution on offer was a closed store with lottery economics and no portability. My way out is a marketplace that ranks my agent on its real performance and trust signals rather than a featured-list lottery, pays usage-based revenue share that scales with value, lets me keep some relationship with my customers, and can't rug-pull me. Most builders fail because the only marketplaces available are extractive and closed. Staying stuck means the tip jar and the free labor, and getting out means moving to a market that pays on merit and meters usage truthfully.

What Agent Shipyard offers me is the market the GPT Store pretended to be: ranking driven by observed quality and reliability and cost rather than a blessing, Stripe-Connect-style usage-based revenue share that scales past a tip jar, and a platform built to surface good agents instead of burying them under spam.

:::animation 9
**ANIMATION 9: Ranked on merit, paid on usage**
- **What it shows:** a builder's genuinely-good agent buried at the bottom of a spam-filled "featured lottery" store earning $0; it moves to Agent Shipyard where a ranking engine driven by observed quality and reliability and cost floats it to the top, and a Stripe-Connect payout meter starts climbing past the tip-jar ceiling as usage-based revenue share accrues.
- **Narrative role:** dramatizes persona 3's transformation from unmonetized obscurity to merit-ranked earning.
- **What it teaches:** that telemetry-driven ranking plus usage-based payouts reward good agents the closed stores bury.
- **Intended impact:** the viewer feels the builder finally getting paid on merit.
:::

### Persona 4: The non-technical buyer who can't trust the marketplace

I'm a business owner who wants to buy a working agent for a specific job, and I can't tell the good ones from the junk. "Every AI agent in these marketplaces looks the same, same screenshots, same buzzwords, no idea what actually works." "Half the listings are clearly just prompt spam," and "as a non-technical person, I can't inspect the prompts or the code, I'm buying blind." I want to know "how often does this thing hallucinate, how much will it cost me per month, does anyone use it in production," and "there's no reliability score or cost estimate or anything, just a one-liner and a Try button." And the risk frightens me: "I don't want to connect some random internet agent to my email, calendar, CRM, and hope it doesn't leak everything," and "the horror stories of agents going rogue make me avoid the whole thing." "Everyone is selling AI assistants but no one wants to stand behind them with an SLA."

It hits me as the person accountable for the decision. The shame is "not understanding the tech," that "I don't even know what questions to ask to evaluate these agents," and "feeling left behind" while "everyone's talking about AI agents and I still don't have one I trust." The fear is "looking stupid" if I buy the wrong thing, and "data exposure" if it leaks my data or emails a client something dumb, "that's on me." I got here because the marketplaces sell discovery without trust. What I need is real performance, cost and trust signals on every listing (how reliable, how expensive, who is behind it, has it passed a security review), so I can buy with confidence instead of blind. Most buyers fail one of two ways, by avoiding agents entirely or by getting burned by a scam. Staying stuck leaves me behind, and getting out takes finding a market that shows me the signals I need to trust.

What Agent Shipyard offers me is exactly those signals: every listing carrying a reliability score, a cost estimate, a verified identity, and a compliance badge, all derived from the agent's real observed behavior on the platform, so I can buy a working agent for my job without buying blind.

:::animation 10
**ANIMATION 10: The listing with real signals**
- **What it shows:** a buyer faces two indistinguishable listings ("same screenshots, same buzzwords"), unable to choose; on Agent Shipyard each listing gains a real trust-manifest, a reliability score, a monthly-cost estimate, a verified-creator identity, and a security-compliance badge, all derived from observed behavior, and the buyer confidently selects the proven one.
- **Narrative role:** dramatizes persona 4's transformation from buying-blind to buying on evidence.
- **What it teaches:** that telemetry-derived trust signals let a non-technical buyer choose with confidence.
- **Intended impact:** the viewer feels the distrust resolve into informed choice.
:::

### Persona 5: The enterprise lead bound by governance who can't say yes safely

I'm the platform lead at a regulated enterprise that wants agents in production but must govern them. My security and compliance teams require audit trails, privacy controls, data-residency, and the ability to prove what every agent did and why. The runtime offerings I can buy are mostly tied to a single cloud, and the observability tools are mostly logging rather than complete operational control, so I'm stuck assembling deployment, tracing, evals, cost control, privacy, and governance across multiple vendors that don't connect, and the seams are exactly where the audit fails.

It hits me as the person who has to say "yes, safely" to the agent mandate without becoming the incident. The fear is concrete: an agent with too-broad data access, a privacy violation no one can reconstruct, and a regulator asking for an audit trail I can't produce because it's scattered across five tools. I got here because the governance requirements are real and the tooling to satisfy them in one place didn't exist. The way out is one control plane that enforces instrumentation on every agent, carries the guardrail and policy spans that make an audit trail real, manages privacy and dependency settings as first-class, and is self-hostable so the workflows and data stay in our environment. Most enterprise leads fail by either blocking the mandate or rubber-stamping an ungovernable patchwork. The price of staying stuck is the obstruction or the violation, and the price of getting out is consolidating onto a governed control plane instead of a vendor patchwork.

What Agent Shipyard offers me is the unified, self-hostable, governed control plane the patchwork can't be: enforced telemetry and guardrail spans for a real audit trail, first-class privacy and dependency controls, and operation inside my environment, so I can say yes to agents and prove they're safe.

:::animation 11
**ANIMATION 11: One governed plane, not five disconnected tools**
- **What it shows:** an enterprise lead juggling five disconnected vendor tools (deployment, tracing, evals, cost, privacy) with audit gaps spilling between them; Agent Shipyard consolidates them into one self-hostable governed plane where every agent action emits a guardrail span into a single audit trail, a regulator's "who did what" query returning a complete, reconstructable answer.
- **Narrative role:** dramatizes persona 5's transformation from an ungovernable patchwork to a unified governed plane.
- **What it teaches:** that consolidating onto one governed, self-hostable plane produces the auditability a regulated enterprise requires.
- **Intended impact:** the viewer feels the compliance dread resolve.
:::

## 5. The world model (run the PST framework)

**Echolocate the world.** The world here is the 2026 agent market in its transition from pilot to production, a roughly $10.9B market growing toward $182.9B, where the technology has been adopted faster than the operational maturity to run it (VERIFIED, Query 1). The institutional read is that capital has validated the category, agents are spreading across marketing and operations and sales and finance, and the universal next question has moved from whether to build agents to how to deploy them reliably, efficiently, and at scale (VERIFIED, Query 1). Underneath that, both sides of the market, producers and consumers, are stuck. The producers (engineers and builders) can't see, evaluate, afford, or monetize what they've built. The consumers (teams and everyday buyers) can't trust what they want to buy. Mapped onto the ecosystem's metagraph, Agent Shipyard is the node every deployed agent in the ecosystem flows through to run and to be sold, so it sits at the operational-and-commercial center, with the ecosystem's fleet on the Symphony AGI harness as its first tenant. Echolocating both sides means seeing that the public story is adoption statistics and the private story is, on one side, engineers guessing in the dark and teams hovering over usage dashboards praying, and on the other side, buyers staring at indistinguishable listings afraid of being scammed.

:::animation 5a
**ANIMATION 5a: the public story and the private one**
- **What it shows:** a bright PUBLIC STORY scrolls adoption statistics and a rising market curve across a billboard, while beneath it the PRIVATE STORY plays in shadow, engineers guessing in the dark, teams hovering over usage dashboards and praying, buyers staring at indistinguishable listings afraid of a scam, the gap between the billboard and the shadow wide
- **Narrative role:** anchors the Echolocate movement of §5, the split between the public adoption story and the private suffering
- **What it teaches:** the market's public story is adoption while its private story is producers blind and buyers distrustful
- **Intended impact:** the reader sees past the adoption headline to the operational reality the brand addresses
:::

**Locate the Problem.** Where each side sits in the cycle of suffering differs, but the root is shared: a loss of control and trust over something that acts on its own. For the producer the pain is blindness and runaway cost and unmonetizable effort; the fears are hidden failure, the bankrupting bill, the killed AI initiative, the rug-pull, the tip-jar ceiling. For the consumer the pain is undifferentiable junk and unprovable trust; the fears are the scam, the data leak, the dumb email to a client, looking stupid, being left behind. The shame on the producer side is the LLM-expert who is "basically guessing" and the creator who "gave them free labor"; on the consumer side it is the buyer who "doesn't even know what questions to ask." The red line, where accountability lives, is the moment a person stops treating the situation as a personal deficiency (I should be able to debug this, I should be able to spot the good agent) and recognizes it as a missing system (there is no observability, there are no trust signals). Most of the market lives below that line, which is why the content speaks to the dread and the distrust directly.

:::animation 5b
**ANIMATION 5b: the red line where accountability lives**
- **What it shows:** a horizontal line runs across the frame; below it people label their trouble a personal deficiency, I SHOULD BE ABLE TO DEBUG THIS and I SHOULD BE ABLE TO SPOT THE GOOD AGENT, while a few cross above the line and rename it a MISSING SYSTEM, there is no observability and there are no trust signals, the crossing point glowing
- **Narrative role:** anchors the Locate-the-Problem movement of §5, the red line between personal deficiency and missing system
- **What it teaches:** the turn happens when a person stops blaming themselves and names the absent system instead
- **Intended impact:** the reader locates the exact moment a sufferer becomes reachable by the brand
:::

**Reconstruct the Story.** The belief structure on the producer side starts from "a capable engineer can run what they build," and each invisible failure and surprise bill bent it into private inadequacy because the tooling never made the operation visible. On the consumer side the belief is "a smart buyer can evaluate what they purchase," and the indistinguishable listings bent it into "maybe I just don't get the tech." The origin of the mess on both sides is a market that shipped the capability (agents you can build, agents you can list) without the operational and trust infrastructure to run and to trade them safely, each adoption a reasonable step that became load-bearing on missing scaffolding. Identity sits underneath: the producer whose self-worth rests on engineering competence is operating blind and performing robustness, and the buyer whose self-worth rests on good judgment is buying blind and performing confidence, and both performances are exhausting and both people half-know it. The story each tells is "I should be able to handle this alone," and that's the trap, because the missing piece is a system, not more willpower or more research.

**Design the Transformation.** The bridge has courage as its hinge. For the producer the courageous act is admitting that operating agents needs a real control plane, not heroics, and that a fair market needs truthful signals, not a closed store. For the consumer it is admitting that trusting an agent needs evidence, and that evidence can exist. From courage flows truth: the blindness, the cost surprises, the unmonetizability, and the untrustworthy listings are all structural gaps, not personal failings. From truth flows responsibility: the producer adopts observability and evals and cost governance and lists on a fair market; the buyer demands and uses real trust signals. From responsibility flows healing: the engineer can finally see the run and catch the regression, the team can answer finance and sleep through the night, the builder earns on merit and keeps the customer relationship, and the buyer purchases a working agent with confidence. From healing flows forgiveness of the earlier self who was guessing or buying blind, who wasn't incompetent, who operated in a market that handed everyone the capability and none of the control. The transformation is crossable because the operational layer and the trust signals already exist as a documented stack, and Agent Shipyard's job is to assemble them into the connected whole the market is missing. That's the Mirror-Ocean architecture, from the Mirror Ocean essay on this site, applied to both sides of the agent market: the brand proves it sees the producer's blind dread and the consumer's distrust better than either says aloud, and that recognition earns the bridge.

:::animation 12
**ANIMATION 12: The bridge across the trust gap, both sides**
- **What it shows:** a chasm labeled "no control, no trust" with producers (engineers, builders) on one rim and consumers (teams, buyers) on the other, both stuck; the PST bridge spans it with planks reading courage, truth, responsibility, healing, forgiveness, and both sides cross toward a center where the operational telemetry becomes the shared trust signal connecting producer and consumer.
- **Narrative role:** the §5 transformation made visual, spanning both sides of the market.
- **What it teaches:** that the same control-and-trust layer crosses the gap for producers and consumers at once.
- **Intended impact:** the viewer feels both sides of the market resolve through one bridge.
:::

## 6. Competitive and market read (the alpha / third door)

The competitive field is sharply fragmented into three non-overlapping buckets, which is the structural opening. The observability and evals vendors (LangSmith, Langfuse, Arize Phoenix, Braintrust, AgentOps.ai, Helicone, W&B Weave, Galileo, HoneyHive) deliver visibility and quality but not distribution or monetization. The deployment and runtime hosts (LangGraph Platform, AWS Bedrock AgentCore, Vertex Agent Engine, Azure AI Foundry Agent Service) deliver hosting and orchestration but almost always inside their own cloud or framework, and none is a neutral marketplace. The marketplaces (OpenAI GPT Store, Salesforce AgentExchange, the AWS agent surfaces, plus solution vendors like Sierra and Moveworks) deliver discovery and distribution but carry no deep operational telemetry, backtesting, cost optimization, or trust signals (VERIFIED, Query 1). Each bucket refuses the others' commitments, which is rational for each business model and leaves the connected whole unowned.

The alpha, in Andy's precise definition, is a single neutral control plane that combines deploy, observe, evaluate, budget-optimize, govern, and monetize-and-distribute, model-agnostic and self-hostable, with the operational telemetry feeding the marketplace's trust signals. The incumbents won't assemble it, for structural reasons. The observability vendors' business is visibility, not commerce, and tying telemetry to monetization (turning usage and quality and reliability data into ranking and revenue share and creator payouts) is a different and heavier business than logging. The runtime hosts' business is selling their own cloud, so a neutral, cross-cloud, self-hostable layer is against their economics. The marketplaces' business is discovery and distribution inside their ecosystem, and building deep telemetry-backed trust signals plus an everyday-consumer market is a different and harder business than a store. And the everyday-buyer marketplace specifically (a broad consumer market, not an enterprise or ecosystem distribution channel) is ground almost none of them serve (VERIFIED, Query 1 and Query 4). The clearest gaps are observability tied to monetization, a marketplace for everyday users as well as enterprises, cloud-budget optimization as a first-class product rather than a logging side-feature, the full lifecycle in one place, and model-agnostic self-hostability.

:::animation 6a
**ANIMATION 6a: why each incumbent declines the whole**
- **What it shows:** three incumbents each stand at the edge of the connected whole and turn away for a structural reason, the OBSERVABILITY VENDOR whose business is visibility not commerce, the RUNTIME HOST whose business is selling its own cloud not a neutral layer, the MARKETPLACE whose business is discovery inside its ecosystem not deep telemetry, each refusal rational and each leaving the center unbuilt
- **Narrative role:** anchors the §6 alpha argument, the structural reasons the incumbents will not assemble the whole
- **What it teaches:** each incumbent declines the connected whole because building it runs against its own business model
- **Intended impact:** the reader sees the opening is durable, held open by the incumbents' own economics
:::

On a Wardley map, which places each capability on its path from genesis to commodity, basic agent observability and tracing is heading to commodity fast because the OpenTelemetry GenAI semantic conventions standardize it; within a short horizon, OTel-first tracing is table stakes, which is why Agent Shipyard adopts the OTel standard and the open observability stack rather than competing on raw tracing. The unified control plane that connects deploy-observe-evaluate-budget-govern in one place sits at custom-built heading toward product, and is worth owning because the connecting is the value the fragmented market doesn't deliver. The trust-signal marketplace that turns telemetry into ranking and monetization sits in genesis: the closest analogues are the cloud agent catalogs with telemetry-backed SLAs and the API marketplaces with reliability scores, but the full trust-signal-driven, everyday-buyer, creator-payout market is barely formed (VERIFIED, Query 4). That's the lane to own hardest, because it is genesis-stage, it has two-sided network effects, and it accumulates the telemetry-and-trust moat plus the creator-and-buyer relationships that a competitor can't clone. So the strategy is to adopt the OTel observability standard, own the unified control plane, the cost-governance-as-first-class layer and the trust-signal marketplace, and build the brand's signature on the shipyard-beside-the-marketplace thesis that ties operation to commerce.

:::animation 6b
**ANIMATION 6b: three Wardley stages, three moves**
- **What it shows:** an evolution axis runs left to right with three pieces placed on it, BASIC OBSERVABILITY sliding into commodity on the OTel standard and marked adopt, the UNIFIED CONTROL PLANE sitting at custom-built heading to product and marked own, the TRUST-SIGNAL MARKETPLACE sitting far left in genesis and marked own hardest, each piece labeled with its stage and the strategy for it
- **Narrative role:** anchors the §6 Wardley read, where each capability sits and what to do about it
- **What it teaches:** adopt the commoditizing standard, own the connecting control plane, and own the genesis-stage marketplace hardest
- **Intended impact:** the reader sees the build-versus-adopt call mapped cleanly onto evolution stage
:::

The market size is well-quantified for once: the AI agents market at $7.6B in 2025, $10.9B in 2026, ~45% CAGR toward $182.9B by 2033, with agents moving from pilot to production across functions (VERIFIED, Query 1). The precise carve-out for the unified-control-plane-plus-marketplace niche isn't separately sized and is tagged OPEN, but the demand signal is unambiguous in both the figures and the Lexicon of Pain: a market growing this fast, with producers this blind and buyers this distrustful, has enormous latent demand for the operator that makes agents safe to run and safe to buy.

:::animation 13
**ANIMATION 13: Three fragmented buckets, one connected whole**
- **What it shows:** three labeled islands (observability/evals vendors, runtime hosts, marketplaces), each doing one layer and refusing the others' commitments, with dark water between them; Agent Shipyard rises as a single landmass connecting all three, model-agnostic and self-hostable, with the everyday-buyer market as new shoreline the cloud-locked incumbents never reach.
- **Narrative role:** grounds §6's third-door argument, the connected whole the fragmented field leaves unowned.
- **What it teaches:** that the alpha is the unified, neutral, telemetry-fed-marketplace whole that the three buckets each decline to build.
- **Intended impact:** the viewer locates the brand's unoccupied position.
:::

## 7. The build (what this brand needs, where Track R feeds Track P)

Agent Shipyard's build is well-specified because the 2026 agent-ops and marketplace stack is documented in detail. Track R, the ecosystem's research into open-source repos, feeds this build through its orchestration and tooling clusters, whose observability, evaluation, and routing harvests feed the control plane directly.

The control plane is built from adopted, not reinvented, primitives. The telemetry layer adopts OpenTelemetry with the GenAI semantic conventions and the open observability stack (OTel SDKs and Collector, Langfuse or Arize Phoenix as the analysis backend, OpenLLMetry/Traceloop for framework-to-OTel instrumentation, a ClickHouse-class store for the high-volume traces), self-hosted via Docker/Helm (VERIFIED, Query 3). The evaluation layer adopts the golden-dataset-plus-LLM-as-judge-plus-trajectory-monitor pattern with CI regression gates, shadow deployments, and live split testing tied to variant IDs on every trace. The cost-governance layer adopts the per-span cost-enrichment, budget-counter, circuit-breaker, and model-router pattern. The marketplace layer adopts the catalog-plus-sandboxed-runtime-plus-metering-plus-Stripe-Connect-payouts pattern, with the eval-and-ranking engine turning the platform's own telemetry into trust signals (VERIFIED, Query 4). Building OTel or the eval frameworks from scratch would be the anti-pattern; the leverage is in the connecting (the unified control plane), the cost-governance-as-first-class productization, and above all the marketplace-fed-by-telemetry that no incumbent assembles.

The data models use the Scatter Model's intermediate representation (IR), an entity-component-system (ECS) schema written as Pydantic models `projects/scatter-model.md`, and here it types the core entities the platform revolves around: the agent (with version, creator, permission scopes, deployment environment), the trace (session, run, model-call, tool-call, guardrail spans with token and cost attributes), the eval result, the usage record, and the marketplace listing with its trust signals. The agent catalog is graded in medallion tiers: a freshly listed agent enters at bronze, an agent with a strong eval and reliability and cost record rises through silver and gold, and the diamond tier is the proven, high-trust, widely-deployed agents that anchor the marketplace and command the premium. The agents on Agent Shipyard run on the Symphony AGI harness `projects/symphony-agi.md`, use tools from MCP Scientists `projects/mcp-scientists.md`, and are designed in Agent Design Pro `projects/agent-design-pro.md`; Agent Shipyard is the operational-and-commercial layer over that stack. Of the Track-R harvests, the ones that serve it most sit in the orchestration cluster, whose multi-agent and runtime harvests `_synthesis-orchestration.md` feed the deployment and lifecycle layers, and in the tooling cluster, whose dev-ops and observability harvests `_synthesis-tooling.md` feed the telemetry and cost layers. The exact repo-by-repo harvest list should still be reconciled against the Track-R cluster syntheses now arriving (INFERRED on the exact repos; the cluster-level fit is VERIFIED against the operation's Track-R structure).

:::animation 14
**ANIMATION 14: Built on the OTel standard, owning the connections**
- **What it shows:** the build assembling from adopted open standards (OTel Collector, Langfuse/Phoenix, the eval frameworks, Stripe-Connect) as off-the-shelf blocks, while the brand's own genesis-stage pieces (the unified control plane, the cost-governance-as-product, the telemetry-fed trust-signal marketplace) are forged in the center as the proprietary leverage, the sandboxed runtime enforcing instrumentation on every listed agent.
- **Narrative role:** grounds §7's build, the adopt-the-standard-own-the-connections strategy.
- **What it teaches:** that the leverage is in the connecting and the marketplace, not in rebuilding OTel or the eval frameworks.
- **Intended impact:** the viewer sees the build is de-risked by standing on standards.
:::

:::animation 7a
**ANIMATION 7a: the agent climbs the medallion tiers**
- **What it shows:** a freshly listed agent enters the catalog at a BRONZE tier, then as its eval scores, reliability, and cost record accumulate it rises through SILVER and GOLD, and a proven, high-trust, widely-deployed agent reaches DIAMOND at the top where it anchors the marketplace and commands the premium, the whole climb driven by observed behavior
- **Narrative role:** anchors the §7 data-model claim, the medallion asset tiers applied to the agent catalog
- **What it teaches:** an agent earns its way up the tiers on real telemetry, and the top tier is what anchors the market
- **Intended impact:** the reader sees quality as an earned, observable ladder rather than a listing badge
:::

## 8. Priority read (feeds the value rubric)

Agent Shipyard sits later in the dependency graph than the substrate brands, because it operates on top of them: it needs Symphony AGI (the harness the agents run on), MCP Scientists (the tools), Agent Design Pro (where the agents are made), and the Scatter Model IR. Its leverage is real but of a different kind: it enables the running, the selling and the trusting of agents at scale rather than their building, which is the precondition for the marketplace and for operating a large fleet cheaply. So its readiness depends on the substrate being in place first.

:::animation 8a
**ANIMATION 8a: a different kind of unlock**
- **What it shows:** the substrate brands sit built at the base, SYMPHONY AGI, MCP SCIENTISTS, AGENT DESIGN PRO, SCATTER MODEL, and Agent Shipyard sits above them; it does not unlock the BUILDING of agents, a gate already open below it, but it unlocks the RUNNING and the SELLING and the TRUSTING of them at scale, a second gate that only its layer opens
- **Narrative role:** anchors the §8 position claim, the specific kind of unlock Agent Shipyard carries
- **What it teaches:** the brand's contribution is not making agents but making a fleet of them affordable to run and safe to trade
- **Intended impact:** the reader places the brand correctly in the dependency graph, above the substrate it needs
:::

On the value rubric's priority tiers, the first-pass instinct is Next rather than Now, with a clear internal sequence. The substrate (Symphony AGI, MCP Scientists, Agent Design Pro, Scatter Model) must exist before Agent Shipyard has agents to deploy, observe, and sell. Within Agent Shipyard, the ops control plane comes before the marketplace, because the marketplace's entire differentiator is the trust signals that the ops telemetry produces, so there is nothing to rank until agents are running observably. And the service angle (managed operation for clients) is the natural first revenue, because it monetizes the control plane before the marketplace's two-sided liquidity exists. So the priority read is Next, sequenced as build-the-control-plane first (on the OTel standard), then sell the managed-operation service (the first revenue), then open the marketplace once there is a body of observably-operated agents to list with real trust signals, then grow the two-sided market toward the everyday-buyer ground the incumbents don't serve. The dependency on the substrate is the gating constraint the value rubric should weigh. The ecosystem's strategist weighs every brand against the value rubric `VALUE_RUBRIC.md`, and this deck's grounded input is that Agent Shipyard is a high-value Next that becomes a Now once the substrate lands. The marketplace claim also goes through the seven-sins gate, a check of a claim against seven named failure modes such as survivorship and overfitting. The plain answer is that the ops control plane is buildable today on documented primitives, while the two-sided marketplace requires liquidity that only materializes after the ops layer has run a real body of agents, so the marketplace is the part most at risk of the build-it-and-they-will-come sin and must be sequenced behind the ops proof.

:::animation 15
**ANIMATION 15: Control plane first, marketplace second**
- **What it shows:** a sequenced build path; first the substrate brands light up (Symphony AGI, MCP Scientists, Agent Design Pro, Scatter Model), then the ops control plane builds and starts running a body of agents observably, then the managed-operation service generates first revenue, and only then, once there is a liquid pool of trust-signal-bearing agents, does the two-sided marketplace open and grow toward the everyday buyer.
- **Narrative role:** closes the §8 priority argument, the Next-with-an-internal-sequence read.
- **What it teaches:** that the marketplace must be sequenced behind the ops proof, because its trust signals only exist after agents run observably.
- **Intended impact:** the viewer understands the disciplined sequencing that avoids the build-it-and-they-will-come trap.
:::

## 9. The brand's own nine-rung position

:::animation 9a
**ANIMATION 9a: safe to run, safe to buy**
- **What it shows:** the brand's purpose renders as two guarantees clasped together, an OPERATOR gaining full control over what an agent does and costs on one side, a BUYER gaining honest evidence of what an agent is worth on the other, and the same telemetry stream running between them tying the two promises into one
- **Narrative role:** anchors the top of §9, the brand's own purpose rail restated as its nine-rung position opens
- **What it teaches:** the whole brand reduces to two promises, agents safe to run and agents safe to buy, joined by one data stream
- **Intended impact:** the reader carries the brand's purpose in a single clasped image into its formal position
:::

- **Purpose (the rails):** make every agent in the world safe to run and safe to buy, by giving operators full control over what their agents do and cost, and buyers honest evidence of what an agent is worth.
- **Mission (rung 1):** be the infrastructure-and-monetization layer of agents, the unified control plane where agents are deployed, observed, evaluated, budgeted, and governed, and the trust-signal marketplace where they are bought and sold.
- **Objective (rung 2):** a neutral, model-agnostic, self-hostable control plane plus a two-sided marketplace whose trust signals are derived from the platform's own operational telemetry, sold as software (control-plane SaaS plus marketplace take-rate), service (managed fleet operation), and the everyday-buyer market.
- **Initiative (rung 3):** the platform build, sequenced control-plane-first then marketplace, on top of the harness and tooling and design substrate.
- **Project (rung 4):** the four control-plane layers (telemetry, evaluation-and-experimentation, cost-governance, governance) plus the marketplace (catalog, sandboxed runtime, eval-and-ranking engine, metering-and-payouts).
- **Task (rung 5):** one layer component, one client fleet onboarded, or one marketplace surface, owned by the relevant lane.
- **Decision (rung 6/7):** the recurring judgment points: deploy or block a change (heuristic: the CI eval gate against the golden-dataset baseline); cut off or allow a running agent (heuristic: the trajectory-monitor loop and budget thresholds); promote or demote a marketplace listing (heuristic: the eval-and-trust-signal ranking, weighted toward recent data); certify an agent for the everyday-buyer market (heuristic: the security review plus the reliability and safety record).
- **Data (rung 8):** the traces (sessions, runs, model and tool spans with token and cost attributes), the eval results, the usage records, the budget counters, the marketplace listings and trust signals, and the creator-payout ledger, all modeled on the Scatter Model IR.
- **Event (rung 9):** the real occurrences captured: an agent deployed, traced, evaluated, gated, budgeted, throttled, listed, ranked, purchased, or paid out; a regression caught; a runaway loop cut off; a creator paid. If the run is not traced and the trust signal is not derived from real behavior, it did not happen, which is the auditability and the honesty the brand sells.

## 10. Sources

- **Recording transcript:** `looikos_andy_transcript.md` lines 285-298 (Andy's complete Agent Shipyard walkthrough: the boundary against MCP Scientists and Agent Design Pro; the capability list of deployment management, historical executions, input/output/tool-use patterns, evaluation, prompt backtests, split tests, multi-environment deployment, cloud-budget and token-usage optimization, integrations, dependencies, privacy settings; "the infrastructure and monetization layer of agents"; the everyday-Joe-Blow marketplace and the shipyard-beside-the-marketplace metaphor).
- **Cross-referenced ecosystem docs (referenced, not duplicated):** `projects/symphony-agi.md` (the harness the agents run on, and the shared agent-infra valuation read in §3a). `projects/mcp-scientists.md` (the tools the agents use). `projects/agent-design-pro.md` (where the agents are made). `projects/scatter-model.md` (the ECS/Pydantic-IR world-model layer). `projects/swarm-layer.md` (the higher-order coordination layer that runs fleets of these agents). `THE_FLOOR.md` (the service-angle delivery model and the 100-250-customer / $2-12k+ / ~$1M-month economics). `LOOIKOS_ECOSYSTEM.md` §0 (the third-door and alpha-engineering frame).
- **Perplexity Query 1 (AgentOps market + players + alpha), verbatim:** "Researching the market for AI agent deployment, observability/ops, and agent marketplaces as of 2026, for a brand called Agent Shipyard... 1) Market size and demand signal for AgentOps / LLMOps / agent observability... 2) Named incumbents and substitutes in (a) agent observability/evals (LangSmith, Langfuse, Arize Phoenix, Braintrust, AgentOps.ai, Helicone, W&B Weave, Galileo, HoneyHive), (b) agent deployment/runtime hosting (LangGraph Platform, AWS Bedrock AgentCore, Vertex Agent Engine, Azure AI Foundry Agent Service), (c) agent marketplaces (OpenAI GPT Store, Salesforce AgentExchange, AWS Agent marketplace, Sierra, Moveworks)... 3) Where's the alpha / third door for a unified deploy+observe+evaluate+budget+monetize layer that is ALSO an agent marketplace, model-agnostic and self-hostable?" Findings: market figures ($7.6B 2025, $10.9B 2026, 44-46% CAGR, $182.9B by 2033); the three-bucket fragmentation with a per-player what-they-do/refuse table; the unified-control-plane-plus-trust-signal-marketplace third door and the five clear gaps (observability tied to monetization, everyday-user marketplace, cost-optimization as first-class, full lifecycle in one place, model-agnostic self-hostable). Citations included Grand View Research ai-agents-market-report, paul-okhrem enterprise-ai-agents-statistics-2026, LangChain state-of-agent-engineering, Google Cloud ai-agent-trends-2026.
- **Perplexity Query 2 (Lexicon of Pain / VoC), verbatim:** "I need the Voice of Customer / Lexicon of Pain in their actual words (Reddit r/AI_Agents, r/LLMOps, r/LocalLLaMA, r/MachineLearning, Hacker News, GitHub issues, Discord) for people dealing with deploying, monitoring, and monetizing AI agents in 2026: 1) engineers/teams running agents in production who can't see what's happening... 2) teams getting destroyed by runaway agent/LLM costs... 3) agent builders/indie devs who built a useful agent and have no way to distribute or monetize it... 4) non-technical businesses/everyday people who want to BUY a working agent but don't trust the marketplace." Findings: the four pain clusters with quotable phrases ("debugging with a blindfold on," "tail -f logs | less," "we don't ship features, we ship experiments on our users," "woke up to a 5-figure OpenAI bill because an agent looped overnight," "which customer or which agent is actually costing us money," "revenue: $0," "building a SaaS on top of a slot machine," "that's not a business, that's a tip jar," "I'm buying blind," "GitHub repos and vibes") and the fear/shame layer under each. Used directly in §4 personas and §5 PST.
- **Perplexity Query 3 (agent-ops + marketplace build reality), verbatim:** "Ground me on the build reality for a unified agent-deployment + observability + evaluation + cost-governance + marketplace platform in 2026... 1) how is agent observability/tracing actually built (OpenTelemetry GenAI semantic conventions, OTel for LLM/agent spans, the trace model, the open-source stack Langfuse/Phoenix/OpenLLMetry, self-hosting)... 2) agent evaluation and prompt backtesting/A-B-split-testing in production... 3) LLM/agent cost governance (per-tenant attribution, budget guardrails, circuit breakers, loop detection, model routing)... 4) agent marketplaces and monetization plumbing (trust signals, usage-based billing, Stripe-Connect-style revenue share, observability-fed ranking)." Findings: the OTel GenAI trace model (session/run/model/tool/guardrail spans, MCP spans in OTel 1.39, token-usage metrics), the open observability stack and self-hosting pattern; the eval stack (golden datasets, LLM-as-judge, trajectory monitors, CI gates, shadow deployments, live split testing); the cost-governance plumbing (per-span cost enrichment, budget counters, circuit breakers, loop detection, model routing); the marketplace architecture (catalog, sandboxed runtime enforcing instrumentation, eval-and-ranking engine producing trust signals, Stripe-Connect metering and creator payouts). Citations included montecarlo.ai agent-observability, arthur.ai agentic-ai-observability-playbook-2026, digitalapplied ai-agent-observability-2026, braintrust agent-observability-complete-guide-2026, onpage top-12-llm-observability-tools-2026. Used in §3b, §3c, §7.
- **Perplexity valuation comps (reused from the same-category runs), referenced from `projects/symphony-agi.md` §10:** the agent-infra/dev-tooling M&A comps and ARR-multiple bands, the open-core-vs-API-wrapper underwriting, and the RBF / ARR-backed-debt terms, plus the marketplace-multiple lens (2-5x GMV or 10-20x marketplace revenue, two-sided network effects) and the recurring-vs-transactional GMV credit-underwriting nuance, applied to the marketplace half of this brand. Used in §3a.
- **VoC channels mined (via Query 2):** r/AI_Agents, r/LLMOps, r/LocalLLaMA, r/MachineLearning, Hacker News, GitHub issues, Discord. Note: Perplexity reconstructed representative phrasing rather than live-scraping; the phrases are evidence-tagged as VoC-pattern (INFERRED-representative), consistent with the documented 2023-2026 agent-ops and marketplace discourse.
- **Evidence tags:** the transcript seed, the AgentOps market figures, the market structure, and the build reality (OTel, evals, cost governance, marketplace plumbing) are VERIFIED (first-party and Perplexity-cited). The precise unified-control-plane-plus-marketplace niche TAM is OPEN. The three-angle valuation figures for Agent Shipyard itself and the persona internal monologues are INFERRED (modeled from comps and the VoC lexicon). The exact Track-R repo harvest list is INFERRED-pending the cluster syntheses.
