Self-containment note (R20): external documents referenced herein are vendored undercanon/as of 2026-07-05. Citations below are the historical record of what this report read at authoring time and are left verbatim; to follow one as a live pointer, resolve the doc undercanon/.
| Field | Value |
|---|---|
| Project | Agent Shipyard |
| Looikos cluster | Infrastructure & Agent Platforms (the infrastructure-and-monetization layer of agents) |
| One-line | The place where agents are deployed, observed, evaluated, budgeted, governed, and monetized, and the marketplace where they are bought and sold, the shipyard beside the great marketplace where the goods come in |
| Status | Concept (depends on the harness and the design and tooling layers it operates on top of) |
1. What it is (the one-paragraph truth)
Agent Shipyard is the infrastructure and monetization layer of agents: the place where agents, once designed and built, are deployed, run, watched, measured, budgeted, governed, and sold. In Andy's words it is where you "deploy and manage the ecosystem of available agents, the marketplace of agents," distinct from where the tools agents use are made (MCP Scientists) and where the agents themselves are designed (Agent Design Pro). Concretely it is two things fused. The first is an agent operations control plane: deployment management across environments, historical execution logging, analysis of input and output and tool-use patterns, evaluation, prompt backtesting and split testing, cloud-budget and token-usage optimization, and the management of integrations, dependencies, and privacy settings, all in one place. The second is a marketplace: a place where available agents are listed, discovered, compared, bought, and deployed, by enterprises and by everyday people, with the operational telemetry from the first half feeding the trust signals of the second half. The name carries the strategy. A shipyard is where vessels are built, fitted, repaired, and launched, and the great marketplaces of history always sat beside the shipyards because that is where the goods come in. Agent Shipyard is positioned at exactly that junction: the operational works where agents are made seaworthy, beside the market where they are traded, with the two reinforcing each other because an agent's observed reliability, cost, and quality become the signals that let a buyer trust it. For the people it serves, it answers a set of miseries that the fragmented tooling market leaves unsolved: engineers flying blind in production, teams terrified of a runaway-cost bill, builders who made something useful and earn nothing from it, and buyers who cannot tell a working agent from a scam.
2. Andy's seed, expanded
Andy's words (verbatim from the recording): "Agent Shipyard is next. So pretty much what we discussed as MCP scientists being for tooling or agentic tooling, so MCPS tools for agents. Agent Shipyard is where we actually deploy and manage the ecosystem of available agents, the marketplace of agents. And so you can look at it as. There's another brand we'll talk about soon called Agent Design Pro. That's where agents are actually made. But here with Agent Shipyard, Agent Shipyard is where you would do things such as managing deployment, looking at historical executions. Obviously we save everything. So you can look at patterns of input, output, tool use. You can see evaluation and run prompt back tests and split tests and manage different deployments across different environments and manage your cloud budget and optimize based on token usage and apply different types of integrations and dependencies and privacy settings. There's a ton of stuff that goes into Agent Shipyards or consider it like the infrastructure layer of agents, the infrastructure and monetization layer of agents. This is also where your everyday Joe Blow, that's where we think like a shipyard. Whereas imagine the massive marketplaces that are all always close to shipyards. There's always going to be a demand for that because this is where the goods are coming in."
Reading between the lines: Andy's seed compresses three claims, each load-bearing.
First, the boundary against its two siblings is stated explicitly and it is the discipline that makes the family of brands coherent rather than one undifferentiated agent-platform. MCP Scientists makes the tools agents call. Agent Design Pro is where agents are made. Agent Shipyard is where made agents are deployed and operated and sold. Keeping these three layers distinct is itself a thesis, because the loudest operational complaint in the 2026 agent market is the tangle of half-overlapping tools where deployment, observability, evaluation, cost, and distribution are split across many vendors and nothing connects. Andy's instinct to draw the line cleanly is the instinct to own the connected whole that the market leaves fragmented.
Second, the enumerated capability list is, read carefully, a precise description of the modern AgentOps control plane, and the order Andy lists them in tracks the actual agent lifecycle. "Managing deployment" and "different deployments across different environments" is the deployment and runtime layer. "Looking at historical executions, we save everything, patterns of input, output, tool use" is the observability and tracing layer, which in 2026 is built OpenTelemetry-first on the GenAI semantic conventions, with sessions and runs and model spans and tool spans and token-usage metrics. "Evaluation and run prompt back tests and split tests" is the evaluation and online-experimentation layer (golden datasets, LLM-as-judge, regression gates in CI, shadow deployments, and live split testing across prompt and model and agent variants). "Manage your cloud budget and optimize based on token usage" is cost governance (per-agent and per-tenant cost attribution, budget guardrails, circuit breakers, loop detection, model routing for cost). "Integrations and dependencies and privacy settings" is the governance layer. Every item Andy names maps onto a real, documented piece of the 2026 stack, which means the brand is not aspirational hand-waving, it is a productization of a real and forming category.
Third, "the infrastructure and monetization layer of agents" plus "your everyday Joe Blow, we think like a shipyard, the marketplaces always close to shipyards, this is where the goods are coming in" is the move that separates Agent Shipyard from every pure-ops tool: it ties the operations to the market. The observability and evaluation data that the ops layer produces is exactly the trust signal a marketplace needs, and no incumbent does this. The observability tools help teams understand behavior but do not turn usage, quality, and reliability data into marketplace ranking, revenue share, or creator payouts; the marketplaces (OpenAI's GPT Store, Salesforce AgentExchange, the cloud agent catalogs) handle discovery and distribution but carry no deep operational telemetry or trust signals. Andy's "shipyard beside the marketplace" is the architectural claim that the operational works and the market belong together, because the operational data is what makes the market trustworthy, and a trustworthy agent market is exactly what the distrustful everyday buyer is waiting for. The "everyday Joe Blow" detail matters too: the existing marketplaces are ecosystem or enterprise distribution channels, not a broad consumer-style market where ordinary people buy a working agent for a specific job, which is a second piece of open ground.
3. The three-angle valuation (the core of a self-standing brand)
Agent Shipyard stands on all three angles with a distinctive shape: its finance angle is the strongest in the family because a marketplace with a take-rate plus usage-metered ops revenue is a payments-and-throughput business that the capital markets value richly; its software angle is a genuine control plane plus a two-sided marketplace; and its service angle is the managed-operation layer for teams that cannot run their own agent fleet.
3a. Finance (credit and capital access)
The category read carries the real numbers this time. The AI agents market is put at roughly $7.6B in 2025 and $10.9B in 2026, growing at a 44-46% CAGR toward $182.9B by 2033. Agent Shipyard sits at the operational-and-commerce center of that market, and the M&A context established in the Symphony AGI deck applies: roughly 90% year-over-year growth in AI M&A deals and a $500M-$5B strategic-valuation band for integrable agentic-AI assets. The comps that matter most here are twofold. On the ops side, the dev-tooling and AI-infra ARR-multiple bands (top-tier AI-infra 12-20x, dev-tools 8-12x, mid-tier 4-8x) and the funded observability vendors (LangSmith, Langfuse, Arize, Braintrust). On the marketplace side, the more interesting comp is the marketplace-business valuation logic: a two-sided market with a take-rate on gross transaction value is valued on marketplace multiples (roughly 2-5x GMV or 10-20x marketplace revenue, in the Stripe-app-store or Shopify-app-store register) rather than pure SaaS multiples, especially when there are real two-sided network effects where more agents attract more buyers and vice versa. Agent Shipyard is the one brand in the family whose finance angle benefits from being read partly as a marketplace, which is a higher-multiple lens than pure infrastructure.
How that converts to credit and capital access has a marketplace-specific nuance. The ops-subscription and managed-operation retainer revenue is contracted and recurring (the good collateral, underwritten like classic SaaS: revenue-based financing at 20-40% of ARR to a 1.2-1.5x cap, ARR-backed debt at 0.3-0.8x ARR, 8-15% plus warrants). The marketplace revenue is the transactional line, and lenders treat it carefully: if a large share of the gross transaction value is recurring automation (agents running ongoing workflows for enterprise buyers), lenders may underwrite a portion of it as quasi-ARR; if it is one-off task purchases, they discount it heavily as transactional. The strategic implication: design the marketplace toward recurring agent subscriptions rather than one-off purchases, because that converts the GMV from heavily-discounted transactional revenue into quasi-recurring revenue with real credit capacity, and it also builds the two-sided retention that the marketplace multiple rewards. The accumulated proprietary state that an acquirer pays the strategic premium for is unusually rich here: the standardized telemetry across every agent that runs on the platform, the trust-and-ranking models trained on that telemetry, the creator and buyer relationships, and the operational track record, none of which a competitor can clone by listing the same agents.
The tri-level market-maker read. Fundamentals: a fast-growing, well-quantified market with Agent Shipyard at its operational-and-commerce center and a defensible two-sided-plus-telemetry position. Technicals: the supply of unified control planes is thin (everyone does one layer), and a credible neutral, model-agnostic, self-hostable one is rarer still, against broad demand, which is a favorable order book. Sentiment: agent operations and governance is the consensus 2026 priority as agents move from pilot to production, with the named risk that the hyperscalers (Bedrock AgentCore, Vertex Agent Engine, Azure AI Foundry) fold ops into their clouds; the hedge is precisely the neutral, model-agnostic, self-hostable, cross-cloud positioning plus the consumer-marketplace ground the hyperscalers do not serve.
3b. Software (the interface stack)
Agent Shipyard's software is a four-layer control plane plus a marketplace, and each layer maps onto a real, documented 2026 architecture.
The telemetry layer is OpenTelemetry-first on the GenAI semantic conventions: a hierarchical trace model of session spans, run spans, model-call spans (carrying gen_ai.request.model and gen_ai.usage token metrics), tool spans (including the MCP spans added in OTel 1.39), retrieval spans, and guardrail spans, collected through an OTel Collector with processors for PII scrubbing, cost enrichment, and per-tenant tagging, and stored in a high-volume backend (ClickHouse-class, the kind Langfuse uses). The evaluation-and-experimentation layer runs golden-dataset regression suites in CI as pre-deploy gates, LLM-as-judge and rule-based evaluators online on sampled live traffic, trajectory monitors that catch the subtle "still kind of works but the workflow changed" regressions, shadow deployments, and live split testing that routes traffic across prompt and model and agent variants with the variant IDs stamped on every trace so metrics slice by variant. The cost-governance layer enriches every model span with a cost_usd from pricing tables, aggregates by tenant and agent and feature, enforces budget guardrails and circuit breakers (max spend per user per day, loop detection via trajectory monitors, max tool calls per session), and routes models for cost (cheap-versus-premium selection driven by the observed quality-versus-cost curve). The governance layer handles integrations, dependencies, privacy settings, and the audit trail.
The marketplace layer is the differentiator and it is built on the telemetry: a catalog of agents as versioned entities with creator identity and required permission scopes; a sandboxed runtime that enforces instrumentation so every listed agent produces standardized telemetry; an eval-and-ranking engine that turns that telemetry into public trust signals (quality scores, reliability SLAs, cost-per-task, usage and retention, compliance badges) and into the ranking that demotes misbehaving or degraded agents quickly; and a metering-and-payouts layer in the Stripe-Connect register (the marketplace is the platform merchant, creators are connected accounts receiving usage-based revenue share, the platform holds funds and handles refunds and retains its fee). The product surfaces and monetization decompose cleanly: the control plane is the SaaS (subscription by agents-under-management, environments, and enterprise features); the ops capabilities are usage-metered where they scale with volume; the marketplace is the take-rate business (a percentage of gross transaction value, in the 10-30% range); and the trust-signal certification can itself be a premium tier. The architectural rule that ties it to the ecosystem: it is model-agnostic and self-hostable, which is the neutral position the cloud-locked incumbents cannot occupy, and it runs the ecosystem's own agents (the harness fleet) as its first and most demanding tenant.
3c. Service (premium-at-accessible boutique delivery)
The service angle is the managed-operation layer: running an agent fleet on behalf of teams that have agents in production and no ability to operate them well. The target operator is the sub-25-employee company or the team that shipped agents and is now flying blind, bleeding cost, and afraid of the next deploy. The premium-quality-at-accessible-pricing model is delivered through the pre-built control plane: because Agent Shipyard already has the telemetry, the evals, the cost guardrails, and the governance built, it can drop a client's agents onto an observable, evaluated, budgeted, governed footing at a fraction of what building that operational stack would cost, and faster.
The retainer economics follow the ecosystem standard: $1-2k accessible at entry, $2-12k+ for the real engagements, structured as a managed-operations retainer (we deploy, observe, evaluate, budget, and govern your agents) plus overage for scale. The 100-250-customer target floors the service angle around $1M/month and scales above. The trust differentiator is the answer to the two deepest operational fears the Lexicon of Pain surfaces: the fear of the runaway-cost bill that "could quietly bankrupt us," answered by real budget guardrails and circuit breakers and per-tenant attribution; and the fear of silent regressions and blind production, answered by the OTel-first tracing and the continuous eval monitors. What gets partnered to the sister affiliate network is the ongoing fleet operation and the long-tail support, run through the shared-floor model with emerging-market senior technical talent operating through the Looikos tools on a franchise/ownership on-ramp, the live transcripts and agent-native systems meaning the floor runs globally. The vertical does not matter; any team running agents in production that needs them observed, evaluated, budgeted, and governed qualifies, and the service angle is also the on-ramp that feeds the marketplace, because a client whose agents are well-operated is a client whose best agents can be listed with real trust signals.
4. The personas (5+, modeled to world-experience depth)
The language here is pulled from the actual Lexicon of Pain mined in the Voice-of-Customer research (Query 2). Agent Shipyard serves both sides of the market, so the personas span the producers (engineers and builders who deploy and want to monetize) and the consumers (teams and everyday people who buy), and the shame textures differ: the producer's shame is "I'm the LLM guy and I'm basically guessing," the consumer's is "I don't even know what questions to ask."
Persona 1: The engineer flying blind in production
I Am the engineer who shipped a "smart" agent and now lives in dread of it. "My agent is basically a black box once it hits production. I have no idea what it's doing step-by-step, just occasional 500s and some vague logs." Multi-step runs "disappear into the void," and when something fails "all I get is tool call failed with zero context." I "can't tell which step of the agent chain is actually breaking," it is "like debugging with a blindfold on." Right now my observability is "tailing logs and grep, that's it." And the regressions are silent: "we tweak one prompt and two weeks later support tickets spike because the agent quietly got dumber in one edge case," with "no alerts, no tests, nothing." I am "terrified of touching the prompts now," everything is "held together with vibes and duct tape," and every release "is a gamble, we don't ship features, we ship experiments on our users."
This impacts me where my professional identity lives. I am "supposed to be the LLM guy, but I'm basically guessing," and the shame is "fake robustness," presenting this thing as AI automation when I know it is duct tape. The fear is hidden failure ("I don't actually know when it's wrong") and getting blamed ("when it screws up in front of a customer, it's on me and I can't even explain what happened"). I got here because shipping the agent was the goal and operating it was an afterthought the tooling did not support. To get out, I need real traces (every thought, tool call, and model response as a nested span with timing and tokens), evaluation gates that catch a regression before it ships, and continuous monitors that alert me when quality drops. Most engineers in my seat fail because they keep adding more printf debugging instead of adopting an observability stack. The cost to stay stuck is the hidden failures and the blame. The cost to get out is admitting that "tail -f logs" is not observability and standing the agent on a real control plane.
What Agent Shipyard offers me is the thing I have been faking: OpenTelemetry-first tracing so I can finally see the whole run, evaluation gates so a prompt change is red or green before it ships, and monitors that tell me when the agent quietly got dumber, so I stop shipping experiments on my users.
Persona 2: The team being bankrupted by runaway cost
I Am the infrastructure person on a team that "woke up to a 5-figure OpenAI bill because an agent looped overnight, no alert, no guardrails, just surprise." Token spend "is a complete black box," and when finance asks "why did this month triple" I have "no answer except the agent talked too much." I "can't answer the most basic question: which customer or which agent is actually costing us money," because "all I get is one giant LLM usage number." There is "no circuit breaker for agents," I "want to say this user can't spend more than five dollars a day and the tooling just doesn't exist," and "our only real guardrail is me hovering over the usage dashboard and praying nothing spikes." Honestly, "I'm more scared of the usage bill than of the bugs at this point," because "every deploy feels like it could quietly bankrupt us if we screw up a loop condition."
This impacts my standing and my sense of responsibility. The shame is being "the infra person" who "can't tell finance where the money is going," and feeling like I am "running a production system with no rate limits and hoping no one notices." The fear is sharper: a dumb prompt mistake that "could literally cost real cash we don't have," and the knowledge that "if this happens again, they're going to shut down the whole AI initiative." I got here because the cost surface of agents is invisible by default and the guardrails are not built in. To get out, I need per-agent and per-tenant cost attribution so I can answer finance, budget guardrails and circuit breakers so a loop cannot bankrupt us, loop detection that cuts off a runaway agent, and model routing so the cheap model handles what it can. Most teams fail because they bolt on a manual hard limit and hope. The cost to stay stuck is the surprise bill and the killed initiative. The cost to get out is treating cost governance as a first-class system instead of a dashboard I stare at.
What Agent Shipyard offers me is cost governance as a real product: every model call cost-attributed by tenant and agent, hard budget guardrails and circuit breakers that stop the loop before it stops us, and routing that keeps the bill rational, so I can finally tell finance exactly where the money goes.
Persona 3: The builder who made a great agent and earns nothing
I Am the indie developer who "built a GPT that users actually love, hundreds of messages, great feedback, revenue: zero." The store "was sold as the App Store for AI and it's basically a random list of spammy bots and SEO bait." "Discovery is nonexistent, unless they bless you in featured, you're invisible," and I am "competing with 10,000 copy-paste SEO GPTs and crypto shills." It is "lottery economics, a tiny handful of creators make real money, everyone else is just adding free content to their platform." Worse, "you don't get any of the user relationship, no email, no way to talk to your customers," and "I'm locked into their store with zero portability." It is "like building a SaaS on top of a slot machine, they can change the odds whenever they want." I built something that "actually saves people hours," but "there's no way to charge per seat, per team, or integrate with their workflows," so "the ceiling is like a hundred dollars a month, that's not a business, that's a tip jar."
This impacts me as someone trying to make a living from real work. The shame is "feeling naïve," that "I bought the App Store for AI dream and ended up giving them free labor," and not having "a real startup," just being "an AI creator on someone else's platform." The fear is "building on shifting sand" where "any day they can change the rules and kill my income." I got here because the only distribution on offer was a closed store with lottery economics and no portability. To get out, I need a marketplace where my agent is ranked on its real performance and trust signals rather than a featured-list lottery, where I get usage-based revenue share that scales with value, where I keep some relationship with my customers, and where I am not rug-pullable. Most builders fail because the only marketplaces available are extractive and closed. The cost to stay stuck is the tip jar and the free labor. The cost to get out is moving to a market that pays on merit and meters honestly.
What Agent Shipyard offers me is the market the GPT Store pretended to be: ranking driven by observed quality and reliability and cost rather than a blessing, Stripe-Connect-style usage-based revenue share that scales past a tip jar, and a platform built to surface good agents instead of burying them under spam.
Persona 4: The non-technical buyer who can't trust the marketplace
I Am a business owner who wants to buy a working agent for a specific job, and I cannot tell the good ones from the junk. "Every AI agent in these marketplaces looks the same, same screenshots, same buzzwords, no idea what actually works." "Half the listings are clearly just prompt spam," and "as a non-technical person, I can't inspect the prompts or the code, I'm buying blind." I want to know "how often does this thing hallucinate, how much will it cost me per month, does anyone use it in production," and "there's no reliability score or cost estimate or anything, just a one-liner and a Try button." And the risk frightens me: "I don't want to connect some random internet agent to my email, calendar, CRM, and hope it doesn't leak everything," and "the horror stories of agents going rogue make me avoid the whole thing." "Everyone is selling AI assistants but no one wants to stand behind them with an SLA."
This impacts me as the person accountable for the decision. The shame is "not understanding the tech," that "I don't even know what questions to ask to evaluate these agents," and "feeling left behind" while "everyone's talking about AI agents and I still don't have one I trust." The fear is "looking stupid" if I buy the wrong thing, and "data exposure" if it leaks my data or emails a client something dumb, "that's on me." I got here because the marketplaces sell discovery without trust. To get out, I need real performance and cost and trust signals on every listing (how reliable, how expensive, who is behind it, has it passed a security review), so I can buy with confidence instead of blind. Most buyers fail by either avoiding agents entirely or getting burned by a scam. The cost to stay stuck is being left behind. The cost to get out is finding a market that shows me the signals I need to trust.
What Agent Shipyard offers me is exactly those signals: every listing carrying a reliability score, a cost estimate, a verified identity, and a compliance badge, all derived from the agent's real observed behavior on the platform, so I can buy a working agent for my job without buying blind.
Persona 5: The enterprise lead bound by governance who can't say yes safely
I Am the platform lead at a regulated enterprise that wants agents in production but must govern them. My security and compliance teams require audit trails, privacy controls, data-residency, and the ability to prove what every agent did and why. The runtime offerings I can buy are mostly tied to a single cloud, and the observability tools are mostly logging rather than complete operational control, so I am stuck assembling deployment, tracing, evals, cost control, privacy, and governance across multiple vendors that do not connect, and the seams are exactly where the audit fails.
This impacts me as the person who has to say "yes, safely" to the agent mandate without becoming the incident. The fear is concrete: an agent with too-broad data access, a privacy violation no one can reconstruct, and a regulator asking for an audit trail I cannot produce because it is scattered across five tools. I got here because the governance requirements are real and the tooling to satisfy them in one place did not exist. To get out, I need one control plane that enforces instrumentation on every agent, carries the guardrail and policy spans that make an audit trail real, manages privacy and dependency settings as first-class, and is self-hostable so the workflows and data stay in our environment. Most enterprise leads fail by either blocking the mandate or rubber-stamping an ungovernable patchwork. The cost to stay stuck is the obstruction or the violation. The cost to get out is consolidating onto a governed control plane instead of a vendor patchwork.
What Agent Shipyard offers me is the unified, self-hostable, governed control plane the patchwork cannot be: enforced telemetry and guardrail spans for a real audit trail, first-class privacy and dependency controls, and operation inside my own environment, so I can say yes to agents and prove they are safe.
5. The world model (run the PST framework)
Echolocate the world. The substrate is the 2026 agent market in its transition from pilot to production, a roughly $10.9B market growing toward $182.9B, where the technology has been adopted faster than the operational maturity to run it. The institutional read: capital has validated the category, agents are spreading across marketing and operations and sales and finance, and the universal next question is no longer whether to build agents but how to deploy them reliably, efficiently, and at scale. Underneath that is a producer-and-consumer market both stuck. The producers (engineers and builders) cannot see, evaluate, afford, or monetize what they have built. The consumers (teams and everyday buyers) cannot trust what they want to buy. The metagraph slice: Agent Shipyard is the node every deployed agent in the ecosystem flows through to run and to be sold, so it sits at the operational-and-commercial center, with the ecosystem's own harness fleet as its first tenant. Echolocating both sides means seeing that the public story is adoption statistics and the private story is, on one side, engineers guessing in the dark and teams hovering over usage dashboards praying, and on the other side, buyers staring at indistinguishable listings afraid of being scammed.
Locate the Problem. The station of the cycle of suffering differs by side but shares a root: a loss of control and trust over something that acts on its own. For the producer the pain is blindness and runaway cost and unmonetizable effort; the fear portfolio is hidden failure, the bankrupting bill, the killed AI initiative, the rug-pull, the tip-jar ceiling. For the consumer the pain is undifferentiable junk and unprovable trust; the fear portfolio is the scam, the data leak, the dumb email to a client, looking stupid, being left behind. The shame on the producer side is the LLM-expert who is "basically guessing" and the creator who "gave them free labor"; on the consumer side it is the buyer who "doesn't even know what questions to ask." The red line, where accountability lives, is the moment a person stops treating the situation as a personal deficiency (I should be able to debug this, I should be able to spot the good agent) and recognizes it as a missing system (there is no observability, there are no trust signals). Most of the market lives below that line, which is why the content speaks to the dread and the distrust directly.
Reconstruct the Story. The belief structure on the producer side starts from "a capable engineer can run what they build," and each invisible failure and surprise bill bent it into private inadequacy because the tooling never made the operation visible. On the consumer side the belief is "a smart buyer can evaluate what they purchase," and the indistinguishable listings bent it into "maybe I just don't get the tech." The origin of the mess on both sides is a market that shipped the capability (agents you can build, agents you can list) without the operational and trust infrastructure to run and to trade them safely, each adoption a reasonable step that became load-bearing on missing scaffolding. The uncomfortable identity layer: the producer whose self-worth rests on engineering competence is operating blind and performing robustness, and the buyer whose self-worth rests on good judgment is buying blind and performing confidence, and both performances are exhausting and both people half-know it. The story each tells is "I should be able to handle this alone," and that is the trap, because the missing piece is a system, not more willpower or more research.
Design the Transformation. The bridge has courage as its hinge. For the producer the courageous act is admitting that operating agents needs a real control plane, not heroics, and that a fair market needs honest signals, not a closed store. For the consumer it is admitting that trusting an agent needs evidence, and that evidence can exist. From courage flows truth: the blindness, the cost surprises, the unmonetizability, and the untrustworthy listings are all structural gaps, not personal failings. From truth flows responsibility: the producer adopts observability and evals and cost governance and lists on a fair market; the buyer demands and uses real trust signals. From responsibility flows healing: the engineer can finally see the run and catch the regression, the team can answer finance and sleep through the night, the builder earns on merit and keeps the customer relationship, and the buyer purchases a working agent with confidence. From healing flows forgiveness of the earlier self who was guessing or buying blind, who was not incompetent, who operated in a market that handed everyone the capability and none of the control. The transformation is crossable because the operational layer and the trust signals already exist as a documented stack, and Agent Shipyard's job is to assemble them into the connected whole the market is missing. This is the Mirror-Ocean architecture applied to both sides of the agent market: the brand proves it sees the producer's blind dread and the consumer's distrust better than either says aloud, and that recognition earns the bridge.
6. Competitive and market read (the alpha / third door)
The competitive field is sharply fragmented into three non-overlapping buckets, which is the structural opening. The observability and evals vendors (LangSmith, Langfuse, Arize Phoenix, Braintrust, AgentOps.ai, Helicone, W&B Weave, Galileo, HoneyHive) deliver visibility and quality but not distribution or monetization. The deployment and runtime hosts (LangGraph Platform, AWS Bedrock AgentCore, Vertex Agent Engine, Azure AI Foundry Agent Service) deliver hosting and orchestration but almost always inside their own cloud or framework, and none is a neutral marketplace. The marketplaces (OpenAI GPT Store, Salesforce AgentExchange, the AWS agent surfaces, plus solution vendors like Sierra and Moveworks) deliver discovery and distribution but carry no deep operational telemetry, backtesting, cost optimization, or trust signals. Each bucket refuses the others' commitments, which is rational for each business model and leaves the connected whole unowned.
The alpha, in Andy's precise definition, is a single neutral control plane that combines deploy, observe, evaluate, budget-optimize, govern, and monetize-and-distribute, model-agnostic and self-hostable, with the operational telemetry feeding the marketplace's trust signals. The reasons the incumbents will not assemble it are structural and worth naming. The observability vendors' business is visibility, not commerce, and tying telemetry to monetization (turning usage and quality and reliability data into ranking and revenue share and creator payouts) is a different and heavier business than logging. The runtime hosts' business is selling their own cloud, so a neutral, cross-cloud, self-hostable layer is against their economics. The marketplaces' business is discovery and distribution inside their ecosystem, and building deep telemetry-backed trust signals plus an everyday-consumer market is a different and harder business than a store. And the everyday-buyer marketplace specifically (a broad consumer market, not an enterprise or ecosystem distribution channel) is ground almost none of them serve. The clearest gaps, stated plainly: observability tied to monetization, a marketplace for everyday users and not just enterprises, cloud-budget optimization as a first-class product rather than a logging side-feature, the full lifecycle in one place, and model-agnostic self-hostability.
The Wardley read. Basic agent observability and tracing is heading to commodity fast because the OpenTelemetry GenAI semantic conventions standardize it; within a short horizon, OTel-first tracing is table stakes, which is why Agent Shipyard adopts the OTel standard and the open observability stack rather than competing on raw tracing. The unified control plane that connects deploy-observe-evaluate-budget-govern in one place sits at custom-built heading toward product, and is worth owning because the connecting is the value the fragmented market does not deliver. The trust-signal marketplace that turns telemetry into ranking and monetization sits in genesis: the closest analogues are the cloud agent catalogs with telemetry-backed SLAs and the API marketplaces with reliability scores, but the full trust-signal-driven, everyday-buyer, creator-payout market is barely formed. That is the lane to own hardest, because it is genesis-stage, it has two-sided network effects, and it accumulates the telemetry-and-trust moat plus the creator-and-buyer relationships that a competitor cannot clone. So the strategy reads: adopt the OTel observability standard, own the unified control plane and the cost-governance-as-first-class layer and the trust-signal marketplace, and signature the brand on the shipyard-beside-the-marketplace thesis that ties operation to commerce.
The market size is well-quantified for once: the AI agents market at $7.6B in 2025, $10.9B in 2026, ~45% CAGR toward $182.9B by 2033, with agents moving from pilot to production across functions. The precise carve-out for the unified-control-plane-plus-marketplace niche is not separately sized, but the demand signal is unambiguous in both the figures and the Lexicon of Pain: a market growing this fast, with producers this blind and buyers this distrustful, has enormous latent demand for the operator that makes agents safe to run and safe to buy.
7. The build (what this brand needs, where Track R feeds Track P)
Agent Shipyard's build is well-specified because the 2026 agent-ops and marketplace stack is documented in detail. The bridge from Track R to Track P here is the orchestration and tooling clusters: the observability, evaluation, and routing harvests feed the control plane directly.
The control plane is built from adopted, not reinvented, primitives. The telemetry layer adopts OpenTelemetry with the GenAI semantic conventions and the open observability stack (OTel SDKs and Collector, Langfuse or Arize Phoenix as the analysis backend, OpenLLMetry/Traceloop for framework-to-OTel instrumentation, a ClickHouse-class store for the high-volume traces), self-hosted via Docker/Helm. The evaluation layer adopts the golden-dataset-plus-LLM-as-judge-plus-trajectory-monitor pattern with CI regression gates, shadow deployments, and live split testing tied to variant IDs on every trace. The cost-governance layer adopts the per-span cost-enrichment, budget-counter, circuit-breaker, and model-router pattern. The marketplace layer adopts the catalog-plus-sandboxed-runtime-plus-metering-plus-Stripe-Connect-payouts pattern, with the eval-and-ranking engine turning the platform's own telemetry into trust signals. Building OTel or the eval frameworks from scratch would be the anti-pattern; the leverage is in the connecting (the unified control plane), the cost-governance-as-first-class productization, and above all the marketplace-fed-by-telemetry that no incumbent assembles.
The data models are the ECS / Pydantic-IR genome (Scatter Model, referenced, see), which here types the core entities the platform revolves around: the agent (with version, creator, permission scopes, deployment environment), the trace (session, run, model-call, tool-call, guardrail spans with token and cost attributes), the eval result, the usage record, and the marketplace listing with its trust signals. The medallion asset tiers apply to the agent catalog: a freshly listed agent enters at bronze, an agent with a strong eval and reliability and cost record rises through silver and gold, and the diamond tier is the proven, high-trust, widely-deployed agents that anchor the marketplace and command the premium. The agents that run on Agent Shipyard run on Symphony AGI (the harness, referenced, see), use tools from MCP Scientists (referenced, see), and are designed in Agent Design Pro (referenced, see); Agent Shipyard is the operational-and-commercial layer over that stack. The Track-R harvests that serve it most are in the orchestration cluster (the multi-agent and runtime harvests in feed the deployment and lifecycle layers) and the tooling cluster (the dev-ops and observability harvests in feed the telemetry and cost layers). The honest note for the lead: the exact repo-by-repo harvest list should be reconciled against the Track-R cluster syntheses now landing.
8. Priority read (feeds the value rubric)
Agent Shipyard sits later in the dependency graph than the substrate brands, because it operates on top of them: it needs Symphony AGI (the harness the agents run on), MCP Scientists (the tools), Agent Design Pro (where the agents are made), and the Scatter Model IR. Its leverage is real but it is leverage of a different kind: it does not unlock the building of agents, it unlocks the running and the selling and the trusting of them at scale, which is the precondition for the marketplace and for operating a large fleet cheaply. So its readiness depends on the substrate being in place first.
The first-pass instinct is Next rather than Now, with a clear internal sequence. The substrate (Symphony AGI, MCP Scientists, Agent Design Pro, Scatter Model) must exist before Agent Shipyard has agents to deploy, observe, and sell. Within Agent Shipyard, the ops control plane comes before the marketplace, because the marketplace's entire differentiator is the trust signals that the ops telemetry produces, so there is nothing to rank until agents are running observably. And the service angle (managed operation for clients) is the natural first revenue, because it monetizes the control plane before the marketplace's two-sided liquidity exists. So the priority read is Next, sequenced as build-the-control-plane first (on the OTel standard), then sell the managed-operation service (the first revenue), then open the marketplace once there is a body of observably-operated agents to list with real trust signals, then grow the two-sided market toward the everyday-buyer ground the incumbents do not serve. The dependency on the substrate is the gating constraint the value rubric should weigh. The strategist reconciles all brands against; this desk's grounded input is that Agent Shipyard is a high-value Next that becomes a Now once the substrate lands, with the seven-sins gate applied to the marketplace claim (the honest answer: the ops control plane is buildable today on documented primitives; the two-sided marketplace requires liquidity that only materializes after the ops layer has run a real body of agents, so the marketplace is the part most at risk of the build-it-and-they-will-come sin and must be sequenced behind the ops proof).