Skip to content
andydataguy

Find the Facts

Infrastructure & agent-platform brand.

Technical Infrastructure~38 min read · 8,892 words
Project
Find the Facts
Looikos cluster
Infrastructure & Agent Platforms (the NLP / text-analysis layer)
One-line
The NLP and text-analysis platform that runs on the knowledge metagraph: extracts facts, entities, claims, statistics, beliefs, and emotional position from text at scale, and is the engine under Easy Insights.
One-line
A data-engineer's dream of NLP over the world's most powerful knowledge metagraph.
Status
Concept (seed is moderate: the NLP/text-analysis platform on the metagraph powering Easy Insights; no standalone repo)

1. What it is (the one-paragraph truth)

Find the Facts is the NLP and text-analysis layer of the Looikos ecosystem: the platform that reads text at scale, extracts meaning from it, and writes that meaning as structured nodes and edges into the knowledge metagraph. Put plainly, it ingests large volumes of text (reports, reviews, transcripts, social posts, articles), extracts the facts, named entities, claims, statistics, and beliefs in that text, and also tags the emotional position of each concept: where it sits on a rich emotional ontology, and whether it pulls a reader toward the cycle of suffering or the cycle of growth, the two loops in Andy's Problem Story Transformation framework (PST) that keep a person stuck or move them forward. Then it writes all of that into a temporal knowledge graph that downstream systems query. Andy's framing is a data-engineer's dream of NLP over the world's most powerful knowledge metagraph, and the role is specific: it's the analysis engine under Easy Insights, the ecosystem's research and intelligence division. The problem it solves is that understanding large bodies of text is still painful, and the existing tools split into two inadequate camps. The conventional NLP and text-analytics platforms and the social and media-intelligence dashboards stop at entity extraction plus coarse sentiment, positive-negative-neutral or at most a fixed list of a few emotions from generic lexicons, and they output scores and word clouds rather than a queryable structure of who believes what, why, and how it changed over time.

The newer LLM tools, the second camp, generate fluent summaries cheaply, but the summaries are ungrounded, with no provenance and no structure. The market research confirms the gap between the two camps: almost nobody is selling deep emotional stance over time in a knowledge metagraph, most platforms run time-series on shallow scores rather than time-aware graphs of beliefs and stance, and conventional sentiment analysis is widely acknowledged as too shallow to capture nuanced emotion or real meaning. Find the Facts sits in that gap with deep entity and claim extraction plus emotional and epistemic position tagging (the emotion a concept carries, and whether a statement is a reported fact, a rumor, or an opinion), grounded in a temporal metagraph with provenance. The market research names that combination as where the LLM moment is pushing demand, because shallow ungrounded summaries are now cheap and deep structured temporal analysis is what sets a product apart.

It serves the ecosystem first, because Easy Insights, the content brands, the quant brands, and Wardley Swarm's grounding all need real text understanding, and second the data and research teams who need to understand text corpora and can't build the pipeline. Find the Facts is concept-stage today with no standalone repo, though the entity-and-emotion analysis it performs is specified in the PST framework's entity-analysis layer (the 130-emotion archive the portfolio keeps in its knowledge graph, ordered on David Hawkins' scale of emotional states from shame at the bottom through courage to peace near the top, and used to tag where each concept sits in the suffering-or-growth cycle), so the brand is the productization of an analysis discipline the ecosystem already defines.

Andy's words (verbatim from the recorded breakdown, where Find the Facts is named inside the Easy Insights walkthrough; lightly de-duplicated, not paraphrased):

Under the hood, Easy Insights is using a tool, a platform I call Find the Facts. And then Find the Facts is like, I mean pretty much take a data engineer mad scientist wet dream about what can you do with NLP, natural language processing and text analysis and then give them a laboratory with the world's most powerful knowledge metagraph data platform. And well, you can imagine what I come up with in this case here. Find the Facts is like that platform under the hood of Easy Insights.

(The transcript above is the canonical seed; the ecosystem overview is a stub and doesn't name Find the Facts. Decompressed to one line, Find the Facts is the NLP and text-analysis platform on the metagraph that powers Easy Insights, a data-engineer's dream of NLP over the world's most powerful knowledge metagraph.)

Reading between the lines. Four things sit compressed in that seed, and the market research confirms each.

The first is the relationship between the metagraph and the analysis, which the phrase "on the metagraph" encodes. The metagraph (WikiDesignCo's world-model) is the substrate, the temporal knowledge graph where facts and concepts and their relationships live with validity windows. Find the Facts is the analysis layer that reads that substrate and, more importantly, writes to it: it's the engine that turns unstructured text into the structured nodes and edges the metagraph holds. So the two do different jobs: the metagraph is the storage and the world-model, and Find the Facts is the NLP that populates and queries it. The research frontier is converging on this architecture: LLMs as extractors and annotators that propose, with the graph as the source of truth that stores and reasons, rather than free-text summarization.

The second is what a data-engineer's dream specifies. It's clean, composable, typed NLP over a temporal knowledge graph, the opposite of today's fragmented setup, where data teams stitch together a cloud NLP API for entities, a separate sentiment model, a vector store for search, and a graph database they have to design and populate themselves, with no unified pipeline. The dream is that the whole pipeline (extract entities, claims, facts, statistics, beliefs, and emotional position, then write typed records into the temporal metagraph with provenance) is one composable platform rather than a custom integration project. The typing comes from Scatter Model, the portfolio's shared typing layer (its intermediate representation, or IR), which is what makes the extracted output structured and trustworthy rather than a free-text dump.

The third is the PST entity-analysis connection, which is the alpha (the hard-to-copy edge) and the part the rest of the market doesn't do. The PST framework specifies that when the platform ingests a chunk of media, it pulls the facts, statistics, beliefs, and claims and also tags where each concept sits in the suffering-or-growth cycle and where it lands on the emotional scale, using the 130-emotion archive and the Hawkins ontology descriptively. Find the Facts is the engine that performs that tagging. This is underserved: existing platforms stop at polarity plus maybe five to eight coarse emotions from domain-agnostic lexicons, and almost nobody operationalizes a rich emotional ontology tied to concepts and tracked over time, even though nuanced emotion (anxiety versus shame versus anger versus hope) is what predicts behavior and reveals narratives. The emotional-position tagging is the signal PST is built to act on and that demographics and shallow sentiment throw away, which is why Find the Facts is the NLP layer that makes the whole PST operating model cheap to run at scale.

The fourth is where it sits in the ecosystem flow, and it's a clean pipeline. Spider Scrape acquires the text from the web, Find the Facts analyzes that text into typed, emotionally-positioned, provenance-tagged structure, WikiDesignCo's metagraph stores it as the world-model, and the consumers read it: Easy Insights as the research and intelligence product (its Sherlock-style corkboard of linked evidence, the deep multi-source analysis with attributed reports), the content brands for the audience and competitive understanding, the quant brands for sentiment and signal, and Wardley Swarm for the evidence its grounded maps cite. Find the Facts is the analysis stage in the acquire-analyze-store-consume flow, which is why it's one of the portfolio's foundational infrastructure brands: the whole grounded-understanding thesis of the ecosystem depends on text being turned into structured meaning, and Find the Facts is what does the turning.

3. The three-angle valuation (the core of a self-standing brand)

3a. Finance (credit and capital access)

The finance read on an NLP and text-analysis platform turns on a structural fact: an analysis pipeline that feeds a customer's research, intelligence, or decision systems is sticky once embedded, because the downstream insights depend on the analysis continuing to arrive in the same structured form, so switching means re-engineering everything that consumes it. That embedding is the credit and valuation foundation, and the category's M&A history shows strategic buyers consistently acquire it.

The economic activity runs on three meters: a usage meter for the analysis volume (the credit-metered analyze-this-text pattern), a seat or subscription meter for the analysts and operators, and an API meter for programmatic consumers. Because Find the Facts is concept-stage with no live revenue, those throughput figures are projections, and the deck holds that. What can be anchored is the quality profile: NLP and intelligence pipelines embed in mission-critical research and decision workflows, so retention is strong once the analysis is load-bearing, and the deeper and more grounded the analysis the harder it is to replace with a commodity sentiment API. That quality of annual recurring revenue (ARR) is what a lender lends against, and the embedded-in-the-workflow stickiness makes the forward revenue forecastable. The capital path is the standard data-and-intelligence-software one: private venture early, with a high-margin premium-intelligence tier (the deep, grounded, emotional-position analysis) lifting revenue per account above commodity NLP.

The M&A and valuation comps are named and post-2020, and they show a clear pattern: strategic buyers in customer-experience, public-relations, and marketing pay for applied NLP and intelligence rather than raw engines. The largest named deal is Brandwatch, acquired by Cision in 2021 for approximately $450M (cash plus shares), combining Cision's media database with Brandwatch's social listening and analytics. The tuck-in pattern is well represented: Lexalytics was acquired by InMoment in 2021 (price undisclosed, a strategic tuck-in to embed text analytics into a customer-experience suite), and MonkeyLearn was acquired by Medallia in 2022 (undisclosed, to give Medallia no-code text classification on feedback data). Primer.ai, the canonical NLP-for-intelligence comp, raised over $150M across rounds. expert.ai is public on the Italian market with a hybrid symbolic-and-ML NLP engine. The adjacency that matters for the metagraph half of the brand is the graph-database market, where Neo4j and TigerGraph have raised large rounds on investor belief in graph-based analytics. The comps show two valuation regimes: sub-$100M strategic tuck-ins for engines (MonkeyLearn, Lexalytics) and $400M-plus for scaled intelligence products (Brandwatch), and Find the Facts is positioned as the latter (an intelligence product, not a raw engine) because the deep grounded analysis is the differentiated value.

Every Looikos brand is valued on three angles (finance, software, and service), each with a $10M floor. Run that floor against those comps and the conclusion holds: $10M is what the service angle alone floors at, and a brand sitting in a category whose comparables run from tuck-ins to $450M, in a combined text-analytics-plus-social-intelligence-plus-graph market triangulated above $20-to-$30B, has a software-angle ceiling well above $10M. The concept-stage discount applies in full: Find the Facts has no ARR, so it's valued today on the thesis, the proven PST entity-analysis discipline, and the category comps, not on a revenue multiple, and the deck projects no fictional ARR.

A market-maker's three-level read (fundamentals, technicals, sentiment) closes it. The fundamentals are the strong embedded-pipeline retention plus the premium-intelligence pricing power, unproven for this specific brand. The technicals are the usage-and-API land-and-expand the category uses. The live sentiment is the most favorable any text-analytics entrant could ask for: cheap LLM summaries raise the value of deep, structured, temporal, emotionally rich, grounded analysis, and enterprise fear of hallucination is driving demand for the grounded, provenance-tracked, source-linked analysis Find the Facts produces. Sentiment is moving toward the analysis the brand does just as the shallow alternative commoditizes, which is the strongest tailwind reading in the category, tempered by the fact that more emotions and graph storage are copyable on their own.

3b. Software (the interface stack)

Software is the core angle for Find the Facts, because the brand is an analysis platform. The product is one analysis core exposed through many surfaces (the hexagonal pattern), and the differentiation lives in the depth of the extraction and the emotional-and-epistemic tagging rather than in commodity entity recognition.

The surfaces map to revenue lines. The analysis platform with its analyst-facing UI is the SaaS subscription surface, and the market research flags it as load-bearing: most graph and NLP vendors ship APIs rather than analyst-friendly tools, so the human-in-the-loop querying and visualization UX (visual graph exploration, natural-language querying such as "show me all negative beliefs about our sustainability efforts in the EU after 2022," and the ability to correct and enrich nodes) is itself a moat rather than a wrapper. The MCP server (a Model Context Protocol endpoint that AI agents can call) is the agent-native surface where other Constellation brands and external agents request analysis on demand, the credit-metered pattern. The CLI and the API support a credit-and-subscription program for programmatic and pipeline consumers. The metagraph read-and-write is the integration surface that makes the brand a primitive rather than a destination: the analysis writes typed nodes and edges into the world-model and queries them back.

The platform decomposes into feature factories, self-contained domains with clean boundaries between them, and five are legible from the seed and the research. The ingestion-and-parse factory takes text from Spider Scrape and other sources into a normalized form. The extraction factory pulls entities, claims, facts, statistics, and beliefs, with the LLM as proposition-extractor and the typed schema as the constraint, so an extraction is who believes what about which entity rather than a bag of keywords. The emotional-and-epistemic-tagging factory adds the 130-emotion-and-Hawkins position tagging plus the epistemic status (whether a statement is a reported fact, a rumor, an opinion, or a counterclaim, with who holds it and how confidently), which the research names as a powerful differentiator. The metagraph-integration factory writes the typed, time-stamped, provenance-tagged records into the temporal world-model. The query-and-analytics factory serves the narrative and driver analytics, the segment-specific stance graphs, and the analyst querying UX. Each follows the custom-modular-composable-harness pattern that Harness V2, the portfolio's shared agent platform, provides.

The market research validates the differentiation most clearly. Emotional-position tagging alone is copyable, since anyone can add more emotions and graph storage, so the alpha is the whole pipeline as one connected primitive: LLM-assisted extraction, emotional and epistemic tagging tied to concepts, the temporal metagraph with provenance, the narrative and driver analytics, and the analyst-friendly querying. Each piece exists somewhere (expert.ai and Ontotext do entity-extraction-into-a-graph, the cloud NLP APIs do entity-level sentiment, the social-intelligence dashboards do coarse emotion labels), but the connected pipeline that turns text into a concept-centric temporal belief-and-emotion graph an analyst can interrogate exists nowhere, and building a full emotional ontology plus the cross-domain stance models plus the querying UX is the multi-year depth incumbents would need to replicate.

The deck builds in two software framings. First, the graph is the product as well as the backend: analysts think in concepts and narratives rather than rows of posts, so the value surfaces when a user can ask how the belief that brand X is overpriced evolved among different segments since 2020 and which events shifted the emotional stance from frustration to resignation, so the concept-centric querying is the product, not the extraction. Second, the strongest adjacent edges in the research (narrative and causal-structure extraction, epistemic-status and provenance, segment-specific stance graphs, task-linked actionable insights) are deepenings of the same pipeline, and the deck treats them as the roadmap above the core rather than as separate products or scattered features. The typed output is the connective tissue: the analysis is typed through Scatter Model's IR and written into WikiDesignCo's metagraph, so the meaning extracted from text becomes world-model structure rather than a dashboard, which applies the portfolio's rule of one authoritative source (the failure it prevents is what the portfolio calls the disconnection) at the point where text becomes knowledge.

3c. Service (premium-at-accessible boutique delivery)

The service angle for Find the Facts is text-intelligence-as-a-service: deliver deep, grounded text and NLP analysis (competitive intelligence, narrative tracking, claims extraction, audience and market understanding) as a retainer, and retain the relationship because the intelligence is ongoing. The delivery moat is the metagraph-grounded depth: analysis that goes beyond keyword counts and shallow sentiment to the concept-centric belief-and-emotion graph an analyst can interrogate, which the market research says today requires human analysts and almost no tool delivers.

The target operator is the portfolio's standard customer resolved to this domain: a master of a craft running a business of fewer than 25 people, who needs to understand large bodies of text and can't build the NLP. These are the researcher or analyst drowning in reports, reviews, and transcripts they can't read at scale, the brand or market-intelligence person who only gets shallow keyword-sentiment dashboards and not real understanding, the founder who needs to understand a market or audience from its text and has no NLP capability, and the strategy or content person who needs to know what is really being said (the claims, the beliefs, the emotional drivers) rather than vibes. They are masters of their domain (the research, the brand, the strategy) who aren't NLP engineers and can't build a text-analysis pipeline, which is that master-of-a-craft profile, and the deep grounded analysis is what lifts the understanding off them.

The engagement shape is the ecosystem standard. An audit at the start locks the scope (which corpora, what depth of analysis, what the deliverable is, what the recurring intelligence cadence is), and the platform quantifies the price against that audit. Premium quality at accessible pricing works because the brand has pre-built the extraction pipeline, the emotional-and-epistemic tagger, and the agent harnesses, so delivering a client's text intelligence is running a proven system with expert oversight rather than building NLP from scratch, which is the compression that lets one operator-architect deliver what a data-science team would. The accessible-product tier sits around the $1-2k/month band and the retainers in the $2-12k+ band, and the service angle floors around $1M/month at the ecosystem-standard 100-to-250 retainer customers.

The commodity analysis beneath the premium engagements (routine sentiment runs, simple entity extraction) gets partnered to the sister affiliate network, and the human operating model that runs the relationship is the portfolio's shared floor, where the team's knowledge lives in a common room and client-facing people work in customer success rather than sales. The service angle has a strong recurring logic specific to this domain: the research names the highest-value use cases as the ones where the stakes are high and the analysis is currently manual (crisis narrative tracking for PR teams, investor-belief mapping for strategy, activist-and-disinformation monitoring), and those are inherently ongoing rather than one-time, because narratives and beliefs and emotional stances evolve continuously, so the intelligence retainer is a continuous belief-and-narrative monitoring relationship rather than a one-off report. That continuity, plus the segment-specific stance graphs that show how different communities' beliefs diverge over time, is what makes the retainer durable and the value compound, and it's the service-side expression of the temporal metagraph that defines the software angle.

4. The personas (5+, modeled to world-experience depth)

This section models six personas in the first person, at world-experience depth, carrying the pain in close-to-real practitioner language. Their phrasing is representative voice: the voice-of-customer research returned constructed-but-realistic phrasings (flagged as such) tightly modeled on how these communities talk and corroborated by analyses of analyst and practitioner complaints, so the phrases are representative rather than documented quotes. The portraits lean toward the negative emotions, because that's where these people live.

P1. The data scientist whose NLP does not understand language

My sentiment model says this review is positive because it sees the word "great" once, but the whole sentence is "great, another update that completely broke the app," and how is that positive in any universe? I have duct-taped together NER and a rule-based sentiment thing and topic models, and it sort of works in the demo, but the second I throw real customer data at it the whole thing collapses, every edge case is a fire drill. We keep shipping dashboards that say eighty-five percent positive sentiment, and then I read the actual comments and customers are clearly furious or sarcastic or begging for help. The model is a vibes detector with brain damage, and I keep adding more rules and the accuracy barely moves.

How it hits my status: I am supposed to be the NLP expert, and I can't get this thing to behave like it actually understands language, so when stakeholders see a rage-post labeled slightly positive I feel like a fraud selling AI that is glorified keyword counting, and I worry a non-technical exec sees a couple of those absurd labels and concludes I don't know what I am doing. How I got here: classical NLP captures word counts, not meaning, and lexicon and classical sentiment models fail on context, irony, domain shift, and nuanced emotion, so the brittleness is structural rather than a personal failing. What it takes to get out: an analysis layer that extracts meaning (the claim, the belief, the actual emotional position) rather than counting words, which is the deep entity-and-emotional-position extraction Find the Facts performs, and it's the difference between a vibes detector and a system that knows great-another-broken-update is fury. Why most stay stuck: the shallow sentiment is the industry default, so its failures read as the inherent limits of NLP rather than a solvable depth problem. The cost of staying stuck is the lying-with-math feeling and the eroding trust in data science. The cost to get out is adopting an analysis layer that extracts meaning so the model stops embarrassing me.

P2. The researcher drowning in text

I have hundreds of interview transcripts and my analysis process is print them out, highlight until my hands hurt, and hope patterns emerge. We have ten thousand open-ended survey responses, leadership wants key themes by Monday, I am one person, and I physically can't read this much text and still think clearly. Every quarter it is the same giant pile of reports and reviews and verbatims, I skim like a maniac, cherry-pick a few quotes, and pray I didn't miss something huge, which isn't a methodology, it is survival. I am drowning in text, and there is probably something really important in there that I will never see because I am too busy firefighting deadlines.

How it hits my status and my life: when I present themes I know deep down they are based on what I happened to read, not the full dataset, and I am scared someone will call that out, and I feel like I am failing as a researcher because the data volume is bigger than my capacity to do real analysis. How I got here: the manual reading and coding of qualitative data doesn't scale, analysts widely report the same pain, and the available tools either give word clouds or require a PhD to use, so I am stuck in copy-paste-into-a-spreadsheet hell. What it takes to get out: an analysis layer that systematically extracts the themes, claims, and emotional drivers across the whole corpus, so the synthesis is grounded in all of it rather than the fraction I could read, which is the analyze-at-scale role Find the Facts is built for. Why most stay stuck: the skim-and-pray method produces something deliverable, so the gap between it and real full-corpus analysis stays hidden until a missed insight surfaces. The cost of staying stuck is the missed critical insights, the burnout, and the fear of being automated away by someone with better tools. The cost to get out is letting a system read the whole corpus so the researcher synthesizes the truth rather than the sample.

P3. The market-intelligence person stuck with shallow dashboards

Our social-listening dashboard shows net sentiment, volume, and a word cloud that says great, love, awesome, and then I click into the actual posts and it is people saying I love how this company never fixes anything, so the tool is a sarcasm amplifier. The CMO keeps asking why do they feel that way, and all I have is net sentiment is up five points and shipping is a trending word, which is embarrassing. These tools are glorified counting machines, they count mentions and positive-versus-negative and keywords, and they don't tell me what people actually believe or what is driving the emotion. Word clouds are the astrology of market research, pretty and vague and useless when I need to explain what is actually going on.

How it hits my status: my job is to provide insight into why people feel what they feel, and the tools give me pretty charts instead of answers, so I am scared leadership thinks we have data and if we miss a shift it is on me, and I worry that when I present these dashboards everyone knows they are shallow and I look like I don't understand the audience. How I got here: the social-intelligence category is dashboard-first and metrics-first, and it runs time-series on shallow scores rather than reasoning about why people feel a certain way or what beliefs drive it, so the why was never in the tool. What it takes to get out: analysis that answers the why, the concept-centric belief-and-emotion graph that shows which beliefs drive the sentiment and how they evolve, which is what Find the Facts produces and what today requires a human analyst. Why most stay stuck: the dashboard looks like data, so the shallowness is socially acceptable until a missed narrative blows up. The cost of staying stuck is the embarrassment in front of the CMO and the replaceability by the next cheap tool. The cost to get out is delivering the why instead of the word cloud.

P4. The founder guessing at a market from its text

I am reading Reddit threads and Discord chats and obscure forum posts trying to understand this market, and I keep ending up with gut feelings instead of anything I would bet the company on. Everyone says talk to your users, but at scale that is thousands of comments and reviews and support tickets, so I am just scrolling and screenshotting and building a narrative in my head, which feels dangerously subjective. I know our customers are telling us exactly what they want in their own words all over the internet, and I can't extract the patterns, so my strategy is read a few spicy posts and guess what the silent majority thinks.

How it hits my status and my life: I am making high-stakes product and market bets based on half-understood conversations, I am afraid someone with better tools or better synthesis sees what I am missing and eats our lunch, and I talk about being customer-obsessed while I can't actually process what customers are saying at scale, which makes me feel like a hypocrite. How I got here: the customer signal is there in the text, inferring customer meaning from reviews and online text is a real strategic problem that basic sentiment and star ratings can't solve, and I have no way to extract the real meaning, so I default to anecdote. What it takes to get out: an analysis layer that turns the audience's own text into a structured read of what they want, fear, and believe, which is the deep audience-understanding Find the Facts produces and the input the whole PST customer-modeling method needs. This persona is the bridge to the PST framework itself, because understanding a market from its text at the level of beliefs and emotional drivers is what PST calls echolocation, modeling a person's whole world rather than a demographic, and Find the Facts is the engine that makes it cheap. Why most stay stuck: the read-a-few-posts method produces a confident-feeling narrative, so the subjectivity stays invisible until a bet built on it fails. The cost of staying stuck is the high-stakes bets on half-understood conversations and the lifelong wonder, if it fails, of whether the users told the truth in plain text and I couldn't read it. The cost to get out is letting a system extract the real patterns so the strategy is grounded in the whole audience.

P5. The comms person who needs the real narrative, not a score

I don't care that sentiment is seventy-two percent positive, I need to know what people are actually saying, what rumors, what misconceptions, what specific things they are mad or excited about. Our monitoring tool gives me X mentions and Y percent positive, and when a crisis hits that is useless, because I need to know what they are accusing us of and what words they are using and what the emotional narrative is. The tool keeps saying overall sentiment stable while on Reddit there is a full-blown conspiracy theory brewing about us, and the system doesn't even have a concept of this is a narrative that can go viral. I end up manually reading the threads anyway because I don't trust the scores, so what is the point of a tool that summarizes vibes and misses the plot.

How it hits my status: if I miss a growing negative narrative because I trusted a sentiment-stable dashboard, that is my name on the incident report, I am supposed to be the person who reads the room and the room is now millions of posts, and I worry the executives think we have it handled because we show them charts when I know the charts aren't capturing the storm brewing underneath. How I got here: the monitoring tools were built around sentiment scores and volume, not around claims and narratives, and they have no first-class concept of a narrative that can go viral or of the specific beliefs and storylines that matter in PR and crisis comms. What it takes to get out: analysis that extracts the actual claims, the narratives, and the emotional drivers and tracks them over time, which is the narrative-and-causal-structure extraction the research names as the strongest adjacent edge and that Find the Facts is built to deliver. Why most stay stuck: the sentiment dashboard is the category standard, so its blindness to narrative is normalized until a narrative it couldn't see becomes a crisis. The cost of staying stuck is the incident report with my name on it and the storm I didn't see. The cost to get out is getting the narrative instead of the score.

P6. The compliance or risk analyst extracting facts from document mountains

I have to extract specific facts and claims from enormous document sets, contracts, filings, disclosures, correspondence, and the stakes are that a missed fact is a compliance failure or a legal exposure. My tools give me keyword search and maybe basic entity recognition, and neither tells me reliably what is claimed, by whom, with what certainty, and citing what evidence, so I read the critical documents by hand and sample the rest and hope the sample was representative. The volume is far beyond what I can read, and the cost of a miss is severe, so I live with a constant low dread that the one fact that mattered was in the ninety percent I didn't get to.

How it hits my status and my life: I am accountable for facts I can't reliably extract at the volume I am given, so the dread is structural, and a single missed claim can become a regulatory or legal event with my name on it. How I got here: the document volume grew, the extraction tools stayed at keyword-and-entity, and the epistemic layer (what is claimed, by whom, how confidently, with what evidence) was never productized, the gap where standard pipelines capture entities and topics but not the claims-and-beliefs-and-evidence structure. What it takes to get out: reliable claim-and-fact extraction with epistemic status and provenance, distinguishing a reported fact from a rumor from an opinion from a counterclaim, which is the epistemic-status-and-provenance edge the research identifies as powerful in compliance and risk and that Find the Facts is built to deliver. Why most stay stuck: keyword search feels like coverage, so the gap between it and reliable claim extraction stays hidden until a missed fact surfaces as an incident. The cost of staying stuck is the constant dread and the genuine regulatory and legal exposure. The cost to get out is reliable epistemic extraction so the facts are found rather than sampled.

5. The world model (run the PST framework)

The six personas share one suffering loop, and modeling it as a single problem-story is what turns the deck from a feature list into PST, which runs in four steps: echolocate, locate the Problem, reconstruct the Story, and design the Transformation.

Echolocate the world. The buyer lives inside a text-understanding ecosystem under a widening gap. On one side is the text, which grows without limit: reviews, transcripts, reports, social posts, filings, forums, every channel producing more than any human can read. On another side is the supply of understanding, which is bifurcated and inadequate: the shallow commodity tools (keyword sentiment, volume, word clouds) that scale but say nothing real, and the deep manual analysis (a human reading and synthesizing) that captures meaning but doesn't scale. On a third side is the new machine option, the LLM, which arrived and reshaped the ground: it makes shallow summaries cheap, which the research says commoditizes the easy capability and therefore raises the value of deep, grounded, structured, temporal analysis, the differentiator. On a fourth side is the rising enterprise fear of hallucination, which is driving demand for grounded, provenance-tracked analysis just as ungrounded summaries become trivial to produce. Read it as an M&A firm reads a target and the leverage is clear: the demand to understand text is universal and growing, the shallow tools are commoditizing, the deep understanding requires human analysts the market can't scale, and the LLM moment is pushing the value toward the grounded temporal depth that's currently unbuilt as a connected product.

Locate the Problem (the cycle of suffering). The pain that arrives is the same for all six: mountains of text I need to understand and can't, at the depth that matters. In response a fear gets installed, and the fear portfolio (the mix of fears the person carries) is specific: the fear of missing the signal (the critical insight in the text I didn't read, the narrative I didn't see, the fact I didn't extract), the fear of the shallow analysis being wrong (the rage-post labeled positive, the sentiment-stable dashboard during a brewing crisis), and the fear of being exposed as not actually understanding the audience or the data I am responsible for. Those fears drive avoidance, which here takes the form of leaning on the inadequate tool or the unscalable manual method rather than solving the depth: the data scientist adds more rules to the vibes detector, the researcher skims and cherry-picks, the intelligence person presents the word cloud, the founder reads a few spicy posts, the comms person trusts the score until they don't, the compliance analyst samples and hopes. The avoidance produces the unfavorable outcome (the misclassification, the missed theme, the embarrassing dashboard, the gut-feel bet, the unseen narrative, the missed fact), and the outcome produces shame, a belief that points at the self instead of at the missing analysis layer: I am a fraud selling glorified keyword counting, I am failing as a researcher, I don't understand my audience, I am a hypocrite about being customer-obsessed. The shame is buried under cope: blame the volume, blame the tools, blame the data, blame the deadline. The red line, the move forbidden, is accountability, because accountability means admitting that decisions and reports were built on text that was never really understood, on shallow scores or unrepresentative samples, and that the gap was tolerated rather than solved. The refusal opens a blind spot, the blind spot produces the next bad action (another rule, another skim, another word cloud, another guess), and the loop closes and compounds, sometimes into a crisis or a failed bet that the un-analyzed text predicted.

Reconstruct the Story. The belief structure under the loop is one of two opposite beliefs that meet at the same trap. For some it's "real understanding requires a human reading it all," so depth and scale are mutually exclusive and the only real analysis is the manual one that can't keep up. For others it's "NLP is just keyword counting," so the machine can only ever give shallow scores and the depth is unreachable by tooling. Both beliefs accept the false dichotomy that you can have scale or meaning but not both, and both keep the person trapped, the first in the unscalable manual grind and the second in the shallow tool. The emotional-experience chain that built it is the information-overwhelm one: the person was rewarded for insight and punished for missing the signal, the text volume outgrew their capacity, and they learned to cope with either heroic manual effort or shallow tooling rather than to demand a deep-and-scalable analysis that didn't seem to exist, so the coping hardened into a belief that the dichotomy is the nature of the problem. The origin layer, where it gets intimate, is the competence-and-comprehension wound: the person's professional identity is built on understanding (the analyst understands the data, the founder understands the market, the comms person reads the room), so admitting they can't actually process the text at the depth and scale required feels like admitting they aren't what their role claims, and the safer move is to keep coping and call the shallow output insight. That's the uncomfortable place most of them run from. On the Hawkins scale used descriptively, the fear and the shame and the comprehension-pride that fuel the loop sit in the destructive band below the courage line.

Design the Transformation. The bridge across hinges on courage. The first step is truth: the scale-or-meaning dichotomy is false, and meaning is extractable and groundable at scale by an analysis layer that extracts entities, claims, beliefs, and emotional position into a temporal graph, so the person can have depth and scale at once and the coping was never necessary. The second is responsibility, owning the reaction rather than the circumstance: the person didn't create the text volume or the shallow tools, but they own whether they keep building decisions on un-analyzed text and calling the shallow output insight. The third is healing, which hurts because it means letting go of either the heroic-manual identity or the NLP-is-shallow resignation and admitting the reports and the dashboards and the bets were built on text that was never really understood, the way the data scientist admits the vibes detector was lying with math and the researcher admits the themes were from the sample. The fourth is forgiveness, releasing the verdict that the comprehension gap is a personal failing, forgiving the misclassifications and the missed themes and the gut-feel bets, and learning from it, which opens the eyes to the new truth that being the person who wields a deep grounded analysis layer is a larger role than being the person who reads heroically or trusts the score. Find the Facts' offer is calibrated to that bridge: the meaning-extraction kills the vibes-detector shame for the data scientist, the analyze-at-scale unblocks the drowning researcher, the why-behind-the-sentiment answers the intelligence person's CMO, the audience-extraction grounds the founder's strategy, the narrative-tracking shows the comms person the storm, and the epistemic extraction finds the compliance analyst's facts. Most of the content lives in the negative band, the misclassification and the drowning and the embarrassing dashboard, because that is where the audience lives, with the deep, grounded, temporal, scalable understanding shown as the reachable other side. That's PST's echolocation method applied to the person who has more text than they can understand and decisions that depend on understanding it.

6. Competitive and market read (the alpha / third door)

The competitive field is layered, crowded at the shallow end and empty at the deep end, and the market research maps it by cluster, by what each cluster refuses to do, and by where the third door sits, the opening no incumbent takes.

Who else does this, and what they won't do. The market sorts into three clusters. The NLP and text-analytics platforms (expert.ai, Lexalytics, MonkeyLearn, Primer.ai, Cohere, Ontotext, John Snow Labs) have strong entity and claim extraction and in some cases knowledge-graph integration, but their sentiment and emotion is crude (polarity plus maybe a few coarse emotions) and not central, and almost none exposes a temporal queryable belief graph: expert.ai does explainable entity extraction and domain taxonomies but emotion-lite sentiment and an internal rather than productized graph, Lexalytics does sentiment and entities OEM'd into customer-experience tools but outputs JSON features not graph nodes, MonkeyLearn (acquired by Medallia in 2022, now part of Medallia's text-analytics stack) does no-code classifiers but tabular labels and no graph, Primer.ai does entity and event extraction for intelligence but is summarization-and-analysis-environments rather than a productized metagraph, Cohere provides embeddings but no out-of-the-box graph and no emotional ontology, and Ontotext and John Snow Labs do entity-and-relation-into-a-graph well but are weak on emotion beyond sentiment. The entity and cloud-NLP layer (Diffbot, AWS Comprehend, Google Cloud Natural Language, Azure) does factual extraction: Diffbot maintains a knowledge graph of billions of entities and trillions of facts but focuses on objective attributes not emotional stance or beliefs, and the cloud APIs do entities and document-or-entity-level polarity sentiment but no native graph, no temporal logic, and no distinction between fact, rumor, opinion, and counterclaim. The social and media-intelligence layer (Brandwatch, Talkwalker, Meltwater, Sprout Social, Sprinklr) is closest to the brand-deep-dive use case but is almost entirely dashboards on shallow analytics: time-series of net sentiment, volume, and word clouds, with at most a fixed list of coarse emotions from generic lexicons, no entity-centric temporal knowledge graph, and no reasoning about why people feel a certain way or what beliefs drive it. Across all three clusters, the market research names the same consistent gaps: deep emotion modeling (everyone stops at polarity plus a few coarse emotions), epistemic structure (nobody captures what is claimed, by whom, with what confidence, citing what evidence, and how it changes over time), temporal knowledge graphs (fact-level temporal reasoning is rare outside research), and cross-corpus concept-centric synthesis (platforms silo by channel rather than unifying into one concept graph).

The third door. Alpha is the thing competitors know about and won't do, and Find the Facts' alpha is the whole pipeline from the software angle as one connected primitive, not any single piece of it. The reason the incumbents won't connect it is structural: the entity-extraction players organize around facts and won't build the rich emotional and epistemic layer, the cloud APIs organize around horizontal infrastructure and won't build the graph or the domain emotion models, the social-intelligence players organize around dashboards and won't build the concept-centric temporal graph, and the emotional ontology, the cross-domain stance models, and the querying UX together would take incumbents years to replicate well. Find the Facts' specific version of the pipeline is distinguished by the PST emotional ontology (the 130-emotion archive and the Hawkins suffering-or-growth positioning, far richer than the coarse-label competitors) and by being native to a temporal metagraph rather than bolted onto a dashboard, which is the combination that exists nowhere.

The pressure-test, and the strongest edges. The market research surfaces the alternative reads, and the deck carries them as the roadmap rather than overclaiming the single feature. The strongest adjacent edges, all deepenings of the same pipeline, are narrative and causal-structure extraction (capturing sequences like brand-raised-prices then customers-felt-betrayed then media-framed-it-as-greed then regulators-intervened as event-and-causal graphs, which is currently a human-analyst task), epistemic status and provenance (who said what, when, with what certainty, citing what evidence, distinguishing fact from rumor from opinion from counterclaim, powerful in disinformation and crisis comms and compliance), segment-specific stance graphs (separate subgraphs for communities showing how their beliefs diverge and converge over time), and task-linked actionable insights (outputting the levers that move the needle, not just the scores).

Wardley evolution and the own-versus-rent call. A Wardley map places each capability on an axis from genesis (new and custom-built) to commodity. Keyword sentiment, basic named-entity recognition, and document-level polarity are commodity, so the brand rents or composes them (the cloud NLP APIs serve where commodity extraction suffices) and never custom-builds them. The rich emotional-and-epistemic tagging, the concept-centric temporal metagraph integration, the narrative-and-causal extraction, and the analyst querying UX are genesis-to-custom: novel, differentiating, load-bearing, the thing competitors won't connect, which is the own-and-build capability where the alpha lives. The embeddings and the base LLM extraction are products to rent (Cohere-class embeddings, the base models) and compose into the pipeline.

Market size and demand signal. The combined addressable space (text analytics plus social and media intelligence plus graph-based knowledge analytics) is triangulated above $20-to-$30B, with the text-analytics market around $6-to-$8B in 2023 growing toward $15-to-$20B-plus by 2028-to-2030 at roughly 18-to-25% CAGR, social listening around $4-to-$7B, and the graph-database market around $3-to-$5B growing toward $10-to-$15B. The demand is revealed, and the LLM moment and enterprise fear of hallucination both amplify it. The category comps confirm the ceiling and the buyer behavior: Brandwatch sold to Cision for roughly $450M, the tuck-ins (MonkeyLearn to Medallia, Lexalytics to InMoment) show strategic buyers paying for applied NLP, and Primer.ai raised over $150M. Demand is proven and rising, the deep connected pipeline is unbuilt, and the LLM moment is pushing value toward the brand's depth.

7. The build (what this brand needs, where Track R feeds Track P)

Find the Facts is concept-stage, so the build section is more provisional than the live brands, but the seed, the PST entity-analysis spec, and the market research pin down the shape.

What it is built from. The extraction stack combines classical NLP and embeddings (the commodity layer, rented or composed) with LLM-based extraction on LangGraph (the LLM as proposition-extractor and stance-annotator and emotion-annotator, the propose-then-verify pattern from the first seed point, where LLMs propose and the graph is the source of truth). The emotional-position tagger is the distinctive component, built on the 130-emotion archive and the Hawkins ontology, tagging where each concept sits on the emotional scale and whether it pulls toward suffering or growth, far richer than the coarse-label competitors. The epistemic tagger captures who claims what, with what certainty, citing what evidence, distinguishing fact from rumor from opinion from counterclaim (a powerful differentiator). The metagraph read-and-write is via Graphiti and Neo4j (the temporal graph with validity windows). The typing is Scatter Model's IR (what makes the output structured and trustworthy). The ingestion comes from Spider Scrape (the acquired text). The querying-and-analytics layer is the narrative-and-driver analytics plus the analyst-facing UX the market research says is itself a moat.

The hexagonal discipline. The design keeps one analysis core with many surfaces. The extraction-and-tagging operations live in a core that never imports a transport, and the analyst UI, the MCP server, the CLI, the API, and the metagraph write are all thin adapters over it. For an analysis layer this is also the defense against what the portfolio calls the disconnection, divergent sources of truth, at the point text becomes knowledge: the typed, emotionally-positioned, provenance-tagged record is the one authoritative representation of an extracted meaning, so the same claim doesn't enter the metagraph three different ways from three different analysis runs.

The data models. The data models are Document, Entity, Claim, Fact, Statistic, Belief, EmotionalPosition, EpistemicStatus, ConceptNode, and StanceEdge, each a typed Pydantic-IR record with temporal validity. The StanceEdge with its emotional position, epistemic status, source, and time span is the load-bearing structure, because it is what turns text into the queryable belief-and-emotion graph.

The agent roster the domain needs. The domain needs four feature factories, each a set of harnesses plus a gateway. The parse-and-ingest factory brings text from Spider Scrape into normalized form. The extraction factory pulls entities, claims, facts, statistics, and beliefs. The emotional-and-epistemic-tagging factory adds the PST position tagging plus the epistemic status, which is the alpha. The metagraph-and-query factory writes the typed records into the temporal world-model and serves the narrative-and-driver analytics and the analyst querying. Each follows the custom-modular-composable-harness pattern the Harness V2 build provides, and the grounded-extraction discipline is shared with Story Factory and Wardley Swarm.

The medallion tiers. The portfolio's medallion tiers (bronze, silver, gold, diamond) map here onto analysis depth: a bronze raw extraction (entities and coarse sentiment), a silver typed extraction (claims and beliefs with emotional position), a gold concept-centric temporal record (the full stance graph with epistemic status and provenance), and a diamond certified narrative-and-driver analysis for a high-stakes consumer (the crisis narrative, the investor-belief map). The provenance the higher tiers carry is what the enterprise-fear-of-hallucination demand is buying.

Where Track R feeds Track P. Track R, the survey of open-source (OSS) repositories that runs beside this brand research (Track P), hasn't started, and for Find the Facts this is a relevant cluster: NLP, entity-extraction, knowledge-graph-construction, and temporal-KG repos are a likely Track-R group. The shape of the need is nameable: Find the Facts will want the best harvested patterns for LLM-assisted entity-and-relation extraction into a graph (the John Snow Labs and Ontotext-class patterns), for temporal knowledge-graph construction and reasoning (the research-frontier temporal-knowledge-graph work), for the emotional and epistemic classification (the domain-tuned emotion-model patterns), and for the analyst-facing graph querying and visualization UX. When the repo decks exist, the value rubric ranks the combined wish-list and the specific capabilities slot in here.

8. Priority read (feeds the value rubric)

Find the Facts is a foundational analysis layer: every text-understanding consumer in the ecosystem depends on it (Easy Insights as its direct product, the content brands for audience understanding, the quant brands for sentiment and signal, Wardley Swarm for the evidence its maps cite, and the PST customer-modeling method itself for echolocation). That makes its leverage high. Its priority is shaped by its position in the pipeline: it is downstream of WikiDesignCo's metagraph (where the analysis lands), Spider Scrape (the text in), and Scatter Model (the IR), and it is the analysis stage between acquisition and storage. On the value rubric's promise-dependency graph it's a high-leverage middle node: many consumers depend on it, and it depends on the metagraph and the IR being real.

Readiness is the real constraint: concept-stage, no standalone repo, so readiness sits below leverage, though the PST entity-analysis discipline being already specified is a stronger readiness signal than a fully unspecified brand carries.

The first-pass tiering runs capability by capability:

  • Next (build and own, gated on the metagraph): the metagraph-grounded extraction-plus-emotional-and-epistemic-tagging pipeline. It's genesis-stage, load-bearing, the alpha competitors won't connect, and high leverage because every text-understanding consumer needs it. It's Next rather than Now because it depends on WikiDesignCo's metagraph being real to write into and on Scatter Model's IR to type the output. The rubric routes it as Powell-VFA, a substrate-shaping analysis layer.
  • Watch (probe before heavy investment): the PST emotional-position tagger and the narrative-and-causal extraction specifically. These are the unique alpha and the strongest adjacent edges, but the research flags that LLM extraction can produce inconsistent results when precision matters and that the emotional ontology and cross-domain stance models are the hard multi-year part, so they route to a probe (validate the emotional-and-epistemic tagging accuracy on real corpora, constrained by the typed schema) before a full commitment. They're genesis-stage and high-potential, the probe profile.
  • Leave (rent and compose, never custom-build): keyword sentiment, basic named-entity recognition, document-level polarity, base embeddings, and base LLM extraction. They're commodity or product, so compose the cloud APIs and the embedding and base-model providers where commodity extraction suffices.

The value rubric's seven-sins gate checks the read against seven biases, each named for a deadly sin. Pride or look-ahead: the read scores the brand concept-stage and the deep pipeline as a bet, and names that any single piece is copyable, so it doesn't score as if the moat already shipped. Envy or survivorship: the failure modes are in the deck (the single-feature-is-copyable risk, the LLM-inconsistency risk, the emotional-ontology-is-hard risk, the metagraph dependency), not just the white-space upside. Gluttony or overfitting: the enthusiasm is capped to the validated whole-pipeline framing and the proven PST entity-analysis discipline, not inflated by the emotional-tagging feature alone. Sloth or transaction-cost: the build friction (the emotional ontology, the cross-domain stance models, the querying UX, the metagraph integration) is named as the gate. Wrath or regime-blindness: the read assumes the 2026 LLM-commoditizes-shallow-summaries regime, which moves value toward the brand's depth, and the enterprise-fear-of-hallucination regime, which drives demand for its grounding. Lust or capacity delusion: Find the Facts is one analysis layer with a probed alpha, not an attempt to win every NLP layer at once. Greed or fat-tail: the tail risk is a social-intelligence incumbent (a Brandwatch) or an entity-extraction player (a Diffbot) deepening into the emotional-and-temporal-graph space, which is why the alpha routes VFA and the moat must be the whole pipeline. The dependency to flag for the portfolio strategist is that Find the Facts' leverage is high (it's the analysis stage every text consumer needs) but it's gated on WikiDesignCo's metagraph and Scatter Model's IR, so it sequences after those, a strong Next in the acquire-analyze-store pipeline (Spider Scrape acquires, Find the Facts analyzes, the metagraph stores), and the smart first move is the metagraph-grounded extraction with the emotional tagging probed before the full narrative-and-causal depth.