Self-containment note (R20): external documents referenced herein are vendored undercanon/as of 2026-07-05. Citations below are the historical record of what this report read at authoring time and are left verbatim; to follow one as a live pointer, resolve the doc undercanon/.
| Field | Value |
|---|---|
| Project | Find the Facts |
| Looikos cluster | Infrastructure & Agent Platforms (the NLP / text-analysis layer) |
| One-line | The NLP and text-analysis platform that runs on the knowledge metagraph: extracts facts, entities, claims, statistics, beliefs, and emotional position from text at scale, and is the engine under Easy Insights. |
| One-line | A data-engineer's dream of NLP over the world's most powerful knowledge metagraph. |
| Status | Concept (seed is moderate: the NLP/text-analysis platform on the metagraph powering Easy Insights; no standalone repo) |
1. What it is (the one-paragraph truth)
Find the Facts is the NLP and text-analysis layer of the Looikos ecosystem: the platform that reads text at scale, extracts meaning from it, and writes that meaning as structured nodes and edges into the knowledge metagraph. The plain version: it ingests large volumes of text (reports, reviews, transcripts, social posts, articles), extracts the facts, named entities, claims, statistics, and beliefs in that text, and crucially also tags the emotional position of each concept, where it sits on a rich emotional ontology and whether it pulls a reader toward the cycle of suffering or the cycle of growth, then it writes all of that into a temporal knowledge graph that downstream systems query. Andy's framing is a data-engineer's dream of NLP over the world's most powerful knowledge metagraph, and the role is specific: it is the analysis engine under Easy Insights, the ecosystem's research and intelligence division. The problem it solves is that understanding large bodies of text remains genuinely painful, and the existing tools split into two inadequate camps. The conventional NLP and text-analytics platforms and the social and media-intelligence dashboards stop at entity extraction plus coarse sentiment, positive-negative-neutral or at most a fixed list of a few emotions from generic lexicons, and they output scores and word clouds rather than a queryable structure of who believes what, why, and how it changed over time.
The newer LLM tools generate fluent summaries cheaply but ungrounded, with no provenance and no structure. The market read confirms this gap precisely: almost nobody is selling deep emotional stance over time in a knowledge metagraph, most platforms run time-series on shallow scores rather than time-aware graphs of beliefs and stance, and conventional sentiment analysis is widely acknowledged as too shallow to capture nuanced emotion or real meaning. Find the Facts sits in that gap: deep entity and claim extraction, plus emotional and epistemic position tagging, grounded in a temporal metagraph with provenance, which the market read names as exactly where the LLM moment is pushing demand because shallow ungrounded summaries are now cheap and deep structured temporal analysis is the differentiator.
The second is what a data-engineer's dream specifies. It is clean, composable, typed NLP over a temporal knowledge graph, the opposite of the fragmented reality the market read describes, where data teams stitch together a cloud NLP API for entities, a separate sentiment model, a vector store for search, and a graph database they have to design and populate themselves, with no unified pipeline. The dream is that the whole pipeline (extract entities, claims, facts, statistics, beliefs, and emotional position, then write typed records into the temporal metagraph with provenance) is one composable platform rather than a custom integration project. The typing is Scatter Model's IR, which is what makes the extracted output structured and trustworthy rather than a free-text dump (referenced, not copied).
The third is the PST entity-analysis connection, which is the alpha and the part the rest of the market does not do. The PST framework specifies that when the platform ingests a chunk of media it does not only pull the facts, statistics, beliefs, and claims; it tags where each concept sits in the suffering-or-growth cycle and where it lands on the emotional scale, using the 130-emotion archive and the Hawkins ontology descriptively. Find the Facts is the engine that performs that tagging. The market read validates this as genuinely underserved: existing platforms stop at polarity plus maybe five to eight coarse emotions from domain-agnostic lexicons, and almost nobody operationalizes a rich emotional ontology tied to concepts and tracked over time, even though nuanced emotion (anxiety versus shame versus anger versus hope) is exactly what predicts behavior and reveals narratives. The emotional-position tagging is the signal PST is built to act on and that demographics and shallow sentiment throw away, which is why Find the Facts is the NLP layer that makes the whole PST operating model cheap to run at scale.
The fourth is where it sits in the ecosystem flow, and it is a clean pipeline. Spider Scrape acquires the text from the web, Find the Facts analyzes that text into typed, emotionally-positioned, provenance-tagged structure, WikiDesignCo's metagraph stores it as the world-model, and the consumers read it: Easy Insights as the research and intelligence product (the Sherlock corkboard, the deep multi-source analysis with attributed reports), the content brands for the audience and competitive understanding, the quant brands for sentiment and signal, and Wardley Swarm for the evidence its grounded maps cite. Find the Facts is the analysis stage in the acquire-analyze-store-consume flow, which is why it is a Category 1 primitive: the whole grounded-understanding thesis of the ecosystem depends on text being turned into structured meaning, and Find the Facts is what does the turning. All siblings (WikiDesignCo, Spider Scrape, Easy Insights, Scatter Model) are referenced, not copied, per the-disconnection.
3. The three-angle valuation (the core of a self-standing brand)
3a. Finance (credit and capital access)
The finance read on an NLP and text-analysis platform turns on a structural fact: an analysis pipeline that feeds a customer's research, intelligence, or decision systems is sticky once embedded, because the downstream insights depend on the analysis continuing to arrive in the same structured form, so switching means re-engineering everything that consumes it. That embedding is the credit and valuation foundation, and the category's M&A history shows it is a category strategic buyers consistently acquire.
The economic activity has three meters. A usage meter for the analysis volume (the credit-metered analyze-this-text pattern), a seat or subscription meter for the analysts and operators, and an API meter for programmatic consumers. Because Find the Facts is concept-stage with no live revenue, those throughput figures are projections, and the deck holds that. What can be anchored is the quality profile: NLP and intelligence pipelines embed in mission-critical research and decision workflows, so retention is strong once the analysis is load-bearing, and the deeper and more grounded the analysis the harder it is to replace with a commodity sentiment API. That ARR quality is what a lender lends against, and the embedded-in-the-workflow stickiness makes the forward revenue forecastable. The capital path is the standard data-and-intelligence-software one: private venture early, with a high-margin premium-intelligence tier (the deep, grounded, emotional-position analysis) lifting revenue per account above commodity NLP.
The M&A and valuation comps are named and post-2020, and they show a clear pattern: strategic buyers in customer-experience, public-relations, and marketing pay for applied NLP and intelligence rather than raw engines. The largest named deal is Brandwatch, acquired by Cision in 2021 for approximately $450M (cash plus shares), combining Cision's media database with Brandwatch's social listening and analytics. The tuck-in pattern is well represented: Lexalytics was acquired by InMoment in 2021 (price undisclosed, a strategic tuck-in to embed text analytics into a customer-experience suite), and MonkeyLearn was acquired by Medallia in 2022 (undisclosed, to give Medallia no-code text classification on feedback data). Primer.ai, the canonical NLP-for-intelligence comp, raised over $150M across rounds. expert.ai is public on the Italian market with a hybrid symbolic-and-ML NLP engine. The adjacency that matters for the metagraph half of the brand is the graph-database market, where Neo4j and TigerGraph have raised large rounds on investor belief in graph-based analytics. The comps show two valuation regimes: sub-$100M strategic tuck-ins for engines (MonkeyLearn, Lexalytics) and $400M-plus for scaled intelligence products (Brandwatch), and Find the Facts is positioned as the latter (an intelligence product, not a raw engine) because the deep grounded analysis is the differentiated value.
Run the $10M floor against those comps and the conclusion holds: $10M is what the service angle alone floors at, and a brand sitting in a category whose comparables run from tuck-ins to $450M, in a combined text-analytics-plus-social-intelligence-plus-graph market triangulated above $20-to-$30B, has a software-angle ceiling well above $10M. The concept-stage discount applies with full honesty: Find the Facts has no ARR, so it is valued today on the thesis, the proven PST entity-analysis discipline, and the category comps, not on a revenue multiple, and the deck projects no fictional ARR.
The market-maker's tri-level read closes it. The fundamentals are the strong embedded-pipeline retention plus the premium-intelligence pricing power, unproven for this specific brand. The technicals are the usage-and-API land-and-expand the category uses. The live sentiment is the most favorable of any text-analytics entrant could ask for, and the market read states it directly: the LLM moment increases the value of this approach, because it makes shallow ungrounded AI summaries cheap and therefore makes deep, structured, temporal, emotionally-rich, grounded analysis the differentiator, and enterprise fear of hallucination is driving demand for exactly the grounded, provenance-tracked, source-linked analysis Find the Facts produces. Sentiment is moving toward the deep grounded temporal analysis the brand is, precisely as the shallow alternative commoditizes, which is the strongest tailwind reading in the category, tempered by the honest note that more-emotions-and-graph-storage is copyable so the moat must be the whole pipeline.
3b. Software (the interface stack)
Software is the core angle for Find the Facts, because the brand is an analysis platform. The product is one analysis core exposed through many surfaces, on the hexagonal core-one-surfaces-many discipline, and the differentiation lives in the depth of the extraction and the emotional-and-epistemic tagging rather than in commodity entity recognition.
The surfaces map to revenue lines. The analysis platform with its analyst-facing UI is the SaaS subscription surface, and the market read flags this as load-bearing: most graph and NLP vendors ship APIs rather than analyst-friendly tools, so the human-in-the-loop querying and visualization UX (visual graph exploration, natural-language querying like show me all negative beliefs about our sustainability efforts in the EU after 2022, and the ability to correct and enrich nodes) is itself a moat rather than a wrapper. The MCP server is the agent-native surface where other Constellation brands and external agents request analysis on demand, the credit-metered pattern. The CLI and the API support a credit-and-subscription program for programmatic and pipeline consumers. The metagraph read-and-write is the integration surface that makes the brand a primitive rather than a destination: the analysis writes typed nodes and edges into the world-model and queries them back.
The platform decomposes into feature factories with clean domain boundaries. Five are legible from the seed and the market read. The ingestion-and-parse factory (taking text from Spider Scrape and other sources into a normalized form). The extraction factory (entities, claims, facts, statistics, and beliefs, with the LLM as proposition-extractor and the typed schema as the constraint, so an extraction is who believes what about which entity rather than a bag of keywords). The emotional-and-epistemic-tagging factory (the 130-emotion-and-Hawkins position tagging plus the epistemic status, whether a statement is a reported fact, a rumor, an opinion, or a counterclaim, with who holds it and how confidently, which the market read names as a powerful differentiator). The metagraph-integration factory (writing the typed, time-stamped, provenance-tagged records into the temporal world-model). The query-and-analytics factory (the narrative and driver analytics, the segment-specific stance graphs, the analyst querying UX). Each is the custom-modular-composable-harness pattern the Harness V2 build provides (referenced from, not copied).
The differentiation is the part the market read validates most clearly and the deck states with the right framing. The alpha is not emotional-position tagging alone, because more-emotions-and-graph-storage is copyable; the alpha is the whole pipeline as one connected primitive: LLM-assisted extraction, emotional and epistemic tagging tied to concepts, the temporal metagraph with provenance, the narrative and driver analytics, and the analyst-friendly querying. Each piece exists somewhere (expert.ai and Ontotext do entity-extraction-into-a-graph, the cloud NLP APIs do entity-level sentiment, the social-intelligence dashboards do coarse emotion labels), but the connected pipeline that turns text into a concept-centric temporal belief-and-emotion graph an analyst can interrogate exists nowhere, and building the robust emotional ontology plus the cross-domain stance models plus the querying UX is exactly the multi-year depth incumbents would need to replicate.
Two honest software framings the deck builds in. First, the graph is the product, not just a backend: the market read is explicit that analysts think in concepts and narratives, not rows of posts, so the value surfaces when a user can ask how the belief that brand X is overpriced evolved among different segments since 2020 and which events shifted the emotional stance from frustration to resignation, which means the concept-centric querying is the product rather than the extraction being the product. Second, the strongest adjacent edges the market read surfaces (narrative and causal-structure extraction, epistemic-status and provenance, segment-specific stance graphs, task-linked actionable insights) are not separate products but deepenings of the same pipeline, and the deck treats them as the roadmap above the core rather than as scattered features. The typed output is the connective tissue: the analysis is typed through Scatter Model's IR and written into WikiDesignCo's metagraph (referenced, not copied), so the meaning extracted from text becomes world-model structure rather than a dashboard, which is the-disconnection discipline applied at the point where text becomes knowledge.
3c. Service (premium-at-accessible boutique delivery)
The service angle for Find the Facts is text-intelligence-as-a-service: deliver deep, grounded text and NLP analysis (competitive intelligence, narrative tracking, claims extraction, audience and market understanding) as a retainer, and retain the relationship because the intelligence is ongoing. The delivery moat is the metagraph-grounded depth: analysis that goes beyond keyword counts and shallow sentiment to the concept-centric belief-and-emotion graph an analyst can interrogate, which is the thing the market read says today requires human analysts and that almost no tool delivers.
The target operator is the Looikos canonical resolved to this domain: the sub-25-employee master-complex who needs to understand large bodies of text and cannot build the NLP. These are the researcher or analyst drowning in reports, reviews, and transcripts they cannot read at scale, the brand or market-intelligence person who only gets shallow keyword-sentiment dashboards and not real understanding, the founder who needs to understand a market or audience from its text and has no NLP capability, and the strategy or content person who needs to know what is really being said (the claims, the beliefs, the emotional drivers) rather than vibes. They are masters of their domain (the research, the brand, the strategy) who are not NLP engineers and cannot build a text-analysis pipeline, which is the master-complex profile, and the deep grounded analysis is what lifts the understanding off them.
The engagement shape is the ecosystem standard. An audit at the start locks the scope (which corpora, what depth of analysis, what the deliverable is, what the recurring intelligence cadence is), and the platform quantifies the price against that audit. Premium quality at accessible pricing works because the brand has pre-built the extraction pipeline, the emotional-and-epistemic tagger, and the agent harnesses, so delivering a client's text intelligence is running a proven system with expert oversight rather than building NLP from scratch, which is the compression that lets one operator-architect deliver what a data-science team would. The accessible-product tier sits around the $1-2k/month band and the retainers in the $2-12k+ band, and the service angle floors around $1M/month at the ecosystem-standard 100-to-250 retainer customers.
The commodity analysis beneath the premium engagements (routine sentiment runs, simple entity extraction) gets partnered to the sister affiliate network, and the human operating model that runs the relationship is the shared-floor customer-success model (referenced from, not copied). The service angle has a strong recurring logic specific to this domain: the market read names the highest-value use cases as the ones where the stakes are high and the analysis is currently manual (crisis narrative tracking for PR teams, investor-belief mapping for strategy, activist-and-disinformation monitoring), and those are inherently ongoing rather than one-time, because narratives and beliefs and emotional stances evolve continuously, so the intelligence retainer is a continuous belief-and-narrative monitoring relationship rather than a one-off report. That continuity, plus the segment-specific stance graphs that show how different communities' beliefs diverge over time, is what makes the retainer durable and the value compound, and it is the service-side expression of the temporal-metagraph that defines the software angle.
4. The personas (5+, modeled to world-experience depth)
Six personas, first person, at world-experience depth, carrying the pain in close-to-real practitioner language. The Lexicon of Pain below is representative voice: the voice-of-customer research returned constructed-but-realistic phrasings (flagged as such) tightly modeled on how these communities talk and corroborated by analyses of analyst and practitioner complaints, so the phrases are tagged representative rather than documented quotes. The bias is toward the negative emotions, because that is where these people live.
P1. The data scientist whose NLP does not understand language
My sentiment model says this review is positive because it sees the word great once, but the whole sentence is great, another update that completely broke the app, and how is that positive in any universe. I have duct-taped together NER and a rule-based sentiment thing and topic models, and it sort of works in the demo, but the second I throw real customer data at it the whole thing collapses, every edge case is a fire drill. We keep shipping dashboards that say eighty-five percent positive sentiment, and then I read the actual comments and customers are clearly furious or sarcastic or begging for help. The model is a vibes detector with brain damage, and I keep adding more rules and the accuracy barely moves.
How it hits my status: I am supposed to be the NLP expert, and I cannot get this thing to behave like it actually understands language, so when stakeholders see a rage-post labeled slightly positive I feel like a fraud selling AI that is glorified keyword counting, and I worry a non-technical exec sees a couple of those absurd labels and concludes I do not know what I am doing. How I got here: classical NLP captures word counts and not meaning, the market read confirms that lexicon and classical sentiment models fail on context, irony, domain shift, and nuanced emotion, so the brittleness is structural rather than a personal failing. What it takes to get out: an analysis layer that extracts meaning (the claim, the belief, the actual emotional position) rather than counting words, which is exactly the deep entity-and-emotional-position extraction Find the Facts performs, and it is the difference between a vibes detector and a system that knows great-another-broken-update is fury. Why most stay stuck: the shallow sentiment is the industry default, so its failures read as the inherent limits of NLP rather than a solvable depth problem. The cost of staying stuck is the lying-with-math feeling and the eroding trust in data science. The cost to get out is adopting an analysis layer that extracts meaning so the model stops embarrassing me.
P2. The researcher drowning in text
I have hundreds of interview transcripts and my analysis process is print them out, highlight until my hands hurt, and hope patterns emerge. We have ten thousand open-ended survey responses, leadership wants key themes by Monday, I am one person, and I physically cannot read this much text and still think clearly. Every quarter it is the same giant pile of reports and reviews and verbatims, I skim like a maniac, cherry-pick a few quotes, and pray I did not miss something huge, which is not a methodology, it is survival. I am drowning in text, and there is probably something really important in there that I will never see because I am too busy firefighting deadlines.
How it hits my status and my life: when I present themes I know deep down they are based on what I happened to read, not the full dataset, and I am scared someone will call that out, and I feel like I am failing as a researcher because the data volume is bigger than my capacity to do real analysis. How I got here: the manual reading and coding of qualitative data genuinely does not scale, the market read confirms this is a widely reported analyst pain, and the available tools either give word clouds or require a PhD to use, so I am stuck in copy-paste-into-a-spreadsheet hell. What it takes to get out: an analysis layer that systematically extracts the themes, claims, and emotional drivers across the whole corpus, so the synthesis is grounded in all of it rather than the fraction I could read, which is the analyze-at-scale role Find the Facts is built for. Why most stay stuck: the skim-and-pray method produces something deliverable, so the gap between it and real full-corpus analysis stays hidden until a missed insight surfaces. The cost of staying stuck is the missed critical insights, the burnout, and the fear of being automated away by someone with better tools. The cost to get out is letting a system read the whole corpus so the researcher synthesizes the truth rather than the sample.
P3. The market-intelligence person stuck with shallow dashboards
Our social-listening dashboard shows net sentiment, volume, and a word cloud that says great, love, awesome, and then I click into the actual posts and it is people saying I love how this company never fixes anything, so the tool is a sarcasm amplifier. The CMO keeps asking why do they feel that way, and all I have is net sentiment is up five points and shipping is a trending word, which is embarrassing. These tools are glorified counting machines, they count mentions and positive-versus-negative and keywords, and they do not tell me what people actually believe or what is driving the emotion. Word clouds are the astrology of market research, pretty and vague and useless when I need to explain what is actually going on.
How it hits my status: my job is to provide insight into why people feel what they feel, and the tools give me pretty charts instead of answers, so I am scared leadership thinks we have data and if we miss a shift it is on me, and I worry that when I present these dashboards everyone knows they are shallow and I look like I do not understand the audience. How I got here: the social-intelligence category is dashboard-first and metrics-first, the market read confirms it runs time-series on shallow scores rather than reasoning about why people feel a certain way or what beliefs drive it, so the why was never in the tool. What it takes to get out: analysis that answers the why, the concept-centric belief-and-emotion graph that shows which beliefs drive the sentiment and how they evolve, which is exactly what Find the Facts produces and what the market read says today requires a human analyst. Why most stay stuck: the dashboard looks like data, so the shallowness is socially acceptable until a missed narrative blows up. The cost of staying stuck is the embarrassment in front of the CMO and the replaceability by the next cheap tool. The cost to get out is delivering the why instead of the word cloud.
P4. The founder guessing at a market from its text
I am reading Reddit threads and Discord chats and obscure forum posts trying to understand this market, and I keep ending up with gut feelings instead of anything I would bet the company on. Everyone says talk to your users, but at scale that is thousands of comments and reviews and support tickets, so I am just scrolling and screenshotting and building a narrative in my head, which feels dangerously subjective. I know our customers are telling us exactly what they want in their own words all over the internet, and I cannot extract the patterns, so my strategy is read a few spicy posts and guess what the silent majority thinks.
How it hits my status and my life: I am making high-stakes product and market bets based on half-understood conversations, I am afraid someone with better tools or better synthesis sees what I am missing and eats our lunch, and I talk about being customer-obsessed while I cannot actually process what customers are saying at scale, which makes me feel like a hypocrite. How I got here: the customer signal is genuinely there in the text, the market read confirms that inferring customer meaning from reviews and online text is a real strategic problem that basic sentiment and star ratings cannot solve, and I have no way to extract the real meaning, so I default to anecdote. What it takes to get out: an analysis layer that turns the audience's own text into a structured read of what they want, fear, and believe, which is exactly the deep audience-understanding Find the Facts produces and the input the whole PST customer-modeling method needs. This persona is the bridge to the PST framework itself, because understanding a market from its text at the level of beliefs and emotional drivers is precisely echolocation, and Find the Facts is the engine that makes echolocation cheap. Why most stay stuck: the read-a-few-posts method produces a confident-feeling narrative, so the subjectivity stays invisible until a bet built on it fails. The cost of staying stuck is the high-stakes bets on half-understood conversations and the lifelong wonder, if it fails, of whether the users told the truth in plain text and I could not read it. The cost to get out is letting a system extract the real patterns so the strategy is grounded in the whole audience.
P5. The comms person who needs the real narrative, not a score
I do not care that sentiment is seventy-two percent positive, I need to know what people are actually saying, what rumors, what misconceptions, what specific things they are mad or excited about. Our monitoring tool gives me X mentions and Y percent positive, and when a crisis hits that is useless, because I need to know what they are accusing us of and what words they are using and what the emotional narrative is. The tool keeps saying overall sentiment stable while on Reddit there is a full-blown conspiracy theory brewing about us, and the system does not even have a concept of this is a narrative that can go viral. I end up manually reading the threads anyway because I do not trust the scores, so what is the point of a tool that summarizes vibes and misses the plot.
How it hits my status: if I miss a growing negative narrative because I trusted a sentiment-stable dashboard, that is my name on the incident report, I am supposed to be the person who reads the room and the room is now millions of posts, and I worry the executives think we have it handled because we show them charts when I know the charts are not capturing the storm brewing underneath. How I got here: the monitoring tools were built around sentiment scores and volume, not around claims and narratives, and the market read confirms they have no first-class concept of a narrative that can go viral or of the specific beliefs and storylines that matter in PR and crisis comms. What it takes to get out: analysis that extracts the actual claims, the narratives, and the emotional drivers and tracks them over time, which is exactly the narrative-and-causal-structure extraction the market read names as the strongest adjacent edge and that Find the Facts is built to deliver. Why most stay stuck: the sentiment dashboard is the category standard, so its blindness to narrative is normalized until a narrative it could not see becomes a crisis. The cost of staying stuck is the incident report with my name on it and the storm I did not see. The cost to get out is getting the narrative instead of the score.
P6. The compliance or risk analyst extracting facts from document mountains
I have to extract specific facts and claims from enormous document sets, contracts, filings, disclosures, correspondence, and the stakes are that a missed fact is a compliance failure or a legal exposure. My tools give me keyword search and maybe basic entity recognition, and neither tells me reliably what is claimed, by whom, with what certainty, and citing what evidence, so I read the critical documents by hand and sample the rest and hope the sample was representative. The volume is far beyond what I can read, and the cost of a miss is severe, so I live with a constant low dread that the one fact that mattered was in the ninety percent I did not get to.
How it hits my status and my life: I am accountable for facts I cannot reliably extract at the volume I am given, so the dread is structural, and a single missed claim can become a regulatory or legal event with my name on it. How I got here: the document volume grew, the extraction tools stayed at keyword-and-entity, and the epistemic layer (what is claimed, by whom, how confidently, with what evidence) was never productized, the exact gap the market read names where standard pipelines capture entities and topics but not the claims-and-beliefs-and-evidence structure. What it takes to get out: reliable claim-and-fact extraction with epistemic status and provenance, distinguishing a reported fact from a rumor from an opinion from a counterclaim, which is exactly the epistemic-status-and-provenance edge the market read identifies as powerful in compliance and risk and that Find the Facts is built to deliver. Why most stay stuck: keyword search feels like coverage, so the gap between it and reliable claim extraction stays hidden until a missed fact surfaces as an incident. The cost of staying stuck is the constant dread and the genuine regulatory and legal exposure. The cost to get out is reliable epistemic extraction so the facts are found rather than sampled.
5. The world model (run the PST framework)
The six personas share one suffering loop, and modeling it as a single problem-story is what turns the deck from a feature list into PST. Echolocate, locate the Problem, reconstruct the Story, design the Transformation.
Echolocate the world. The buyer lives inside a text-understanding ecosystem under a widening gap. On one side is the text, which grows without limit: reviews, transcripts, reports, social posts, filings, forums, every channel producing more than any human can read. On another side is the supply of understanding, which is bifurcated and inadequate: the shallow commodity tools (keyword sentiment, volume, word clouds) that scale but say nothing real, and the deep manual analysis (a human reading and synthesizing) that captures meaning but does not scale. On a third side is the new machine option, the LLM, which arrived and reshaped the ground: it makes shallow summaries cheap, which the market read says commoditizes the easy capability and therefore raises the value of deep, grounded, structured, temporal analysis, the differentiator. On a fourth side is the rising enterprise fear of hallucination, which is driving demand for grounded, provenance-tracked analysis precisely as ungrounded summaries become trivial to produce. Read it as an M&A firm reads a target and the leverage is clear: the demand to understand text is universal and growing, the shallow tools are commoditizing, the deep understanding requires human analysts the market cannot scale, and the LLM moment is pushing the value toward exactly the grounded temporal depth that is currently unbuilt as a connected product.
Locate the Problem (the cycle of suffering). The pain that arrives is the same for all six: mountains of text I need to understand and cannot, at the depth that matters. In response a fear gets installed, and the fear portfolio is specific. The fear of missing the signal (the critical insight in the text I did not read, the narrative I did not see, the fact I did not extract), the fear of the shallow analysis being wrong (the rage-post labeled positive, the sentiment-stable dashboard during a brewing crisis), and the fear of being exposed as not actually understanding the audience or the data I am responsible for. Those fears drive avoidance, which here takes the form of leaning on the inadequate tool or the unscalable manual method rather than solving the depth: the data scientist adds more rules to the vibes detector, the researcher skims and cherry-picks, the intelligence person presents the word cloud, the founder reads a few spicy posts, the comms person trusts the score until they do not, the compliance analyst samples and hopes. The avoidance produces the unfavorable outcome (the misclassification, the missed theme, the embarrassing dashboard, the gut-feel bet, the unseen narrative, the missed fact), and the outcome produces shame, the belief not I lack a deep analysis layer but I am a fraud selling glorified keyword counting, I am failing as a researcher, I do not understand my audience, I am a hypocrite about being customer-obsessed. The shame is buried under cope: blame the volume, blame the tools, blame the data, blame the deadline. The red line, the move forbidden, is accountability, because accountability means admitting that decisions and reports were built on text that was never really understood, on shallow scores or unrepresentative samples, and that the gap was tolerated rather than solved. The refusal opens a blind spot, the blind spot produces the next bad action (another rule, another skim, another word cloud, another guess), and the loop closes and compounds, sometimes into a crisis or a failed bet that the un-analyzed text predicted.
Reconstruct the Story. The belief structure under the loop is one of two opposite beliefs that meet at the same trap. For some it is real understanding requires a human reading it all, so depth and scale are mutually exclusive and the only honest analysis is the manual one that cannot keep up. For others it is NLP is just keyword counting, so the machine can only ever give shallow scores and the depth is unreachable by tooling. Both beliefs accept the false dichotomy that you can have scale or meaning but not both, and both keep the person trapped, the first in the unscalable manual grind and the second in the shallow tool. The emotional-experience chain that built it is the information-overwhelm one: the person was rewarded for insight and punished for missing the signal, the text volume outgrew their capacity, and they learned to cope with either heroic manual effort or shallow tooling rather than to demand a deep-and-scalable analysis that did not seem to exist, so the coping hardened into a belief that the dichotomy is the nature of the problem. The origin layer, where it gets intimate, is the competence-and-comprehension wound: the person's professional identity is built on understanding (the analyst understands the data, the founder understands the market, the comms person reads the room), so admitting they cannot actually process the text at the depth and scale required feels like admitting they are not what their role claims, and the safer move is to keep coping and call the shallow output insight. That is the uncomfortable place most of them run from. On the Hawkins scale used descriptively, the fear and the shame and the comprehension-pride that fuel the loop sit in the destructive band below the courage line.
Design the Transformation. The bridge across hinges on courage. The first step is truth, and the uncomfortable truth is that the scale-or-meaning dichotomy is false, that meaning is extractable and groundable at scale by an analysis layer that extracts entities, claims, beliefs, and emotional position into a temporal graph, so the person can have depth and scale at once and the coping was never necessary. The second is responsibility, owning the reaction rather than the circumstance: the person did not create the text volume or the shallow tools, but they own whether they keep building decisions on un-analyzed text and calling the shallow output insight. The third is healing, which hurts because it means letting go of either the heroic-manual identity or the NLP-is-shallow resignation and admitting the reports and the dashboards and the bets were built on text that was never really understood, the way the data scientist admits the vibes detector was lying with math and the researcher admits the themes were from the sample. The fourth is forgiveness, releasing the verdict that the comprehension gap is a personal failing, forgiving the misclassifications and the missed themes and the gut-feel bets, and learning from it, which opens the eyes to the new truth that being the person who wields a deep grounded analysis layer is a larger role than being the person who reads heroically or trusts the score. Find the Facts' offer is calibrated to that bridge: the meaning-extraction kills the vibes-detector shame for the data scientist, the analyze-at-scale unblocks the drowning researcher, the why-behind-the-sentiment answers the intelligence person's CMO, the audience-extraction grounds the founder's strategy, the narrative-tracking shows the comms person the storm, and the epistemic extraction finds the compliance analyst's facts. Most of the content lives in the negative band, the misclassification and the drowning and the embarrassing dashboard, because that is where the audience lives, with the deep, grounded, temporal, scalable understanding shown as the reachable other side. That is the Echolocation architecture applied to the person who has more text than they can understand and decisions that depend on understanding it.
6. Competitive and market read (the alpha / third door)
The competitive field is layered and crowded at the shallow end and empty at the deep end, which the market read makes clear. Map it by cluster, by what each refuses, and by where the third door is.
Who else does this, and what they will not do. Three clusters. The NLP and text-analytics platforms (expert.ai, Lexalytics, MonkeyLearn, Primer.ai, Cohere, Ontotext, John Snow Labs) have strong entity and claim extraction and in some cases knowledge-graph integration, but their sentiment and emotion is crude (polarity plus maybe a few coarse emotions) and not central, and almost none exposes a temporal queryable belief graph: expert.ai does explainable entity extraction and domain taxonomies but emotion-lite sentiment and an internal rather than productized graph, Lexalytics does sentiment and entities OEM'd into customer-experience tools but outputs JSON features not graph nodes, MonkeyLearn (acquired by Medallia in 2022, now part of Medallia's text-analytics stack) does no-code classifiers but tabular labels and no graph, Primer.ai does entity and event extraction for intelligence but is summarization-and-analysis-environments rather than a productized metagraph, Cohere provides embeddings but no out-of-the-box graph and no emotional ontology, and Ontotext and John Snow Labs do entity-and-relation-into-a-graph well but are weak on emotion beyond sentiment. The entity and cloud-NLP layer (Diffbot, AWS Comprehend, Google Cloud Natural Language, Azure) does factual extraction: Diffbot maintains a knowledge graph of billions of entities and trillions of facts but focuses on objective attributes not emotional stance or beliefs, and the cloud APIs do entities and document-or-entity-level polarity sentiment but no native graph, no temporal logic, and no distinction between fact, rumor, opinion, and counterclaim. The social and media-intelligence layer (Brandwatch, Talkwalker, Meltwater, Sprout Social, Sprinklr) is closest to the brand-deep-dive use case but is almost entirely dashboards on shallow analytics: time-series of net sentiment, volume, and word clouds, with at most a fixed list of coarse emotions from generic lexicons, no entity-centric temporal knowledge graph, and no reasoning about why people feel a certain way or what beliefs drive it. Across all three clusters, the consistent gaps the market read names are the same: deep emotion modeling (everyone stops at polarity plus a few coarse emotions), epistemic structure (nobody captures what is claimed, by whom, with what confidence, citing what evidence, and how it changes over time), temporal knowledge graphs (fact-level temporal reasoning is rare outside research), and cross-corpus concept-centric synthesis (platforms silo by channel rather than unifying into one concept graph).
The third door. Alpha is the thing competitors know about and will not do, and Find the Facts' alpha is the whole pipeline as one connected primitive, not any single piece. The market read is explicit that emotional-position tagging alone is not the moat because more-emotions-and-graph-storage is copyable, and that the defensible differentiation is the whole pipeline: LLM-assisted extraction, emotional and epistemic tagging tied to concepts, the temporal metagraph with provenance, the narrative and driver analytics, and the analyst-friendly querying and visualization. The reason the incumbents will not connect it is structural: the entity-extraction players organize around facts and will not build the rich emotional and epistemic layer, the cloud APIs organize around horizontal infrastructure and will not build the graph or the domain emotion models, the social-intelligence players organize around dashboards and will not build the concept-centric temporal graph, and building the robust emotional ontology plus the cross-domain stance models plus the querying UX is exactly the multi-year depth the market read says would take incumbents years to replicate well. Find the Facts' specific version of the pipeline is distinguished by the PST emotional ontology (the 130-emotion archive and the Hawkins suffering-or-growth positioning, far richer than the coarse-label competitors) and by being native to a temporal metagraph rather than bolted onto a dashboard, which is the combination that exists nowhere.
The honest pressure-test, and the strongest edges. The market read surfaces the alternative reads and the deck carries them as the roadmap rather than overclaiming the single feature. The strongest adjacent edges, all deepenings of the same pipeline, are narrative and causal-structure extraction (capturing sequences like brand-raised-prices then customers-felt-betrayed then media-framed-it-as-greed then regulators-intervened as event-and-causal graphs, which is currently a human-analyst task), epistemic status and provenance (who said what, when, with what certainty, citing what evidence, distinguishing fact from rumor from opinion from counterclaim, powerful in disinformation and crisis comms and compliance), segment-specific stance graphs (separate subgraphs for communities showing how their beliefs diverge and converge over time), and task-linked actionable insights (outputting the levers that move the needle, not just the scores). The deck treats these as the depth roadmap above the core, which is exactly the market read's synthesis that the whole pipeline is the moat and the graph is the product.
Wardley evolution and the own-versus-rent call. Keyword sentiment, basic named-entity recognition, and document-level polarity are commodity, rent or compose (the cloud NLP APIs serve where commodity extraction suffices), never custom-build. The rich emotional-and-epistemic tagging, the concept-centric temporal metagraph integration, the narrative-and-causal extraction, and the analyst querying UX are genesis-to-custom: novel, differentiating, load-bearing, the thing competitors will not connect, which is the own-and-build capability where the alpha lives. The embeddings and the base LLM extraction are product, rent (Cohere-class embeddings, the base models), compose into the pipeline.
Market size and demand signal. The combined addressable space (text analytics plus social and media intelligence plus graph-based knowledge analytics) is triangulated above $20-to-$30B, with the text-analytics market around $6-to-$8B in 2023 growing toward $15-to-$20B-plus by 2028-to-2030 at roughly 18-to-25% CAGR, social listening around $4-to-$7B, and the graph-database market around $3-to-$5B growing toward $10-to-$15B. The demand is revealed and the LLM moment amplifies it: shallow summaries are now cheap, which commoditizes the easy capability and makes deep grounded structured temporal analysis the differentiator, and enterprise fear of hallucination is driving demand for exactly the grounded provenance-tracked analysis the brand produces. The category comps confirm the ceiling and the buyer behavior: Brandwatch sold to Cision for roughly $450M, the tuck-ins (MonkeyLearn to Medallia, Lexalytics to InMoment) show strategic buyers paying for applied NLP, and Primer.ai raised over $150M. Demand is proven and rising, the deep connected pipeline is unbuilt, and the LLM moment is pushing value toward exactly the brand's depth, which is the wave it rides, tempered by the honest note that the moat must be the whole pipeline because any single piece is copyable.
7. The build (what this brand needs, where Track R feeds Track P)
Find the Facts is concept-stage, so the build section is more provisional than the live brands, but the seed, the PST entity-analysis spec, and the market read pin down the shape.
What it is built from. The extraction stack combines classical NLP and embeddings (the commodity layer, rented or composed) with LLM-based extraction on LangGraph (the LLM as proposition-extractor and stance-annotator and emotion-annotator, which the market read names as the pattern where LLMs propose and the graph is the source of truth). The emotional-position tagger is the distinctive component, built on the 130-emotion archive and the Hawkins ontology, tagging where each concept sits on the emotional scale and whether it pulls toward suffering or growth, far richer than the coarse-label competitors. The epistemic tagger captures who claims what, with what certainty, citing what evidence, distinguishing fact from rumor from opinion from counterclaim (the market read's powerful differentiator). The metagraph read-and-write is via Graphiti and Neo4j (the temporal graph with validity windows). The typing is Scatter Model's IR (what makes the output structured and trustworthy). The ingestion comes from Spider Scrape (the acquired text). The querying-and-analytics layer is the narrative-and-driver analytics plus the analyst-facing UX the market read says is itself a moat.
The hexagonal discipline. One analysis core, surfaces many. The extraction-and-tagging operations live in a core that never imports a transport, and the analyst UI, the MCP server, the CLI, the API, and the metagraph write are all thin adapters over it. For an analysis layer this is also the defense against the Disconnection at the point text becomes knowledge: the typed, emotionally-positioned, provenance-tagged record is the one authoritative representation of an extracted meaning, so the same claim does not enter the metagraph three different ways from three different analysis runs, which is the divergent-sources-of-truth failure the-disconnection names, prevented at the analysis boundary.
The data models. Document, Entity, Claim, Fact, Statistic, Belief, EmotionalPosition, EpistemicStatus, ConceptNode, and StanceEdge, each a typed Pydantic-IR record with temporal validity. The StanceEdge with its emotional position, epistemic status, source, and time span is the load-bearing structure, because it is what turns text into the queryable belief-and-emotion graph.
The agent roster the domain needs. Four feature factories, each a set of harnesses plus a gateway. The parse-and-ingest factory (text from Spider Scrape into normalized form). The extraction factory (entities, claims, facts, statistics, beliefs). The emotional-and-epistemic-tagging factory (the PST position tagging plus the epistemic status, the alpha). The metagraph-and-query factory (writing the typed records into the temporal world-model and serving the narrative-and-driver analytics and the analyst querying). Each is the custom-modular-composable-harness pattern the Harness V2 build provides (referenced from, not copied), and the grounded-extraction discipline is shared with Story Factory and Wardley Swarm (referenced, not copied).
The medallion tiers. Applied to analysis depth: a bronze raw extraction (entities and coarse sentiment), a silver typed extraction (claims and beliefs with emotional position), a gold concept-centric temporal record (the full stance graph with epistemic status and provenance), and a diamond certified narrative-and-driver analysis for a high-stakes consumer (the crisis narrative, the investor-belief map). The provenance the higher tiers carry is what the enterprise-fear-of-hallucination demand is buying.
Where Track R feeds Track P. Track R has not started, and for Find the Facts this is a relevant cluster: NLP, entity-extraction, knowledge-graph-construction, and temporal-KG repos are a likely Track-R group. The shape of the need is nameable: Find the Facts will want the best harvested patterns for LLM-assisted entity-and-relation extraction into a graph (the John Snow Labs and Ontotext-class patterns), for temporal knowledge-graph construction and reasoning (the research-frontier TKG work the market read cites), for the emotional and epistemic classification (the domain-tuned emotion-model patterns), and for the analyst-facing graph querying and visualization UX. When the repo decks exist at, the value rubric ranks the combined wish-list and the specific capabilities slot in here.
8. Priority read (feeds the value rubric)
Find the Facts is a foundational analysis layer: every text-understanding consumer in the ecosystem depends on it (Easy Insights as its direct product, the content brands for audience understanding, the quant brands for sentiment and signal, Wardley Swarm for the evidence its maps cite, and the PST customer-modeling method itself for echolocation). That makes its leverage high. Its priority is shaped by its position in the pipeline: it is downstream of WikiDesignCo's metagraph (where the analysis lands), Spider Scrape (the text in), and Scatter Model (the IR), and it is the analysis stage between acquisition and storage. On the promise-dependency graph it is a high-leverage middle node: many consumers depend on it, and it depends on the metagraph and the IR being real.
Readiness is the honest constraint: concept-stage, no standalone repo, so readiness sits below leverage, though the PST entity-analysis discipline being already specified is a stronger readiness signal than a fully unspecified brand carries.
The first-pass tiering, capability by capability:
- Next (build and own, gated on the metagraph): the metagraph-grounded extraction-plus-emotional-and-epistemic-tagging pipeline. It is genesis-stage, load-bearing, the alpha competitors will not connect, and high leverage because every text-understanding consumer needs it. It is Next rather than Now because it depends on WikiDesignCo's metagraph being real to write into and on Scatter Model's IR to type the output. Routes Powell-VFA (substrate-shaping analysis layer).
- Watch (probe before heavy investment): the PST emotional-position tagger and the narrative-and-causal extraction specifically. These are the unique alpha and the strongest adjacent edges, but the market read flags that LLM extraction can produce inconsistent results when precision matters and that the emotional ontology and cross-domain stance models are the hard multi-year part, so they route to a probe (validate the emotional-and-epistemic tagging accuracy on real corpora, constrained by the typed schema) before a full commitment. Genesis, high-potential, the probe profile.
- Leave (rent and compose, never custom-build): keyword sentiment, basic named-entity recognition, document-level polarity, base embeddings, and base LLM extraction. Commodity or product. Compose the cloud APIs and the embedding and base-model providers where commodity extraction suffices.
Run the seven-sins gate. Pride or look-ahead: the read scores the brand concept-stage and the deep pipeline as a bet, and names that any single piece is copyable, so it does not score as if the moat already shipped. Envy or survivorship: the failure modes are in the deck (the single-feature-is-copyable risk, the LLM-inconsistency risk, the emotional-ontology-is-hard risk, the metagraph dependency), not just the white-space upside. Gluttony or overfitting: the enthusiasm is capped to the validated whole-pipeline framing and the proven PST entity-analysis discipline, not inflated by the emotional-tagging feature alone. Sloth or transaction-cost: the build friction (the emotional ontology, the cross-domain stance models, the querying UX, the metagraph integration) is named as the gate. Wrath or regime-blindness: the read assumes the 2026 LLM-commoditizes-shallow-summaries regime, which moves value toward the brand's depth, and the enterprise-fear-of-hallucination regime, which drives demand for its grounding. Lust or capacity delusion: Find the Facts is one analysis layer with a probed alpha, not an attempt to win every NLP layer at once. Greed or fat-tail: the tail risk is a social-intelligence incumbent (a Brandwatch) or an entity-extraction player (a Diffbot) deepening into the emotional-and-temporal-graph space, which is why the alpha routes VFA and the moat must be the whole pipeline. The dependency to flag for the strategist: Find the Facts' leverage is high (it is the analysis stage every text consumer needs) but it is gated on WikiDesignCo's metagraph and Scatter Model's IR, so it sequences after those, a strong Next in the acquire-analyze-store pipeline (Spider Scrape acquires, Find the Facts analyzes, the metagraph stores), and the smart first move is the metagraph-grounded extraction with the emotional tagging probed before the full narrative-and-causal depth.