# The Source-First Manifesto

> Five theses on source-first intelligence: why market and competitor monitoring should start with the sources, not with keywords. With evidence and objections.

- Author: Dr. Karsten Richter
- Publisher: Picasi GmbH
- Topic: What is Source-First Intelligence?
- Last updated: 2026-09-23
- Canonical URL: https://source-first-intelligence.com/en/manifest/
- Language: en


Source-First Intelligence is market and competitor monitoring that starts with the sources: you decide whose publications to follow and analyze their own content, instead of searching the open web for keywords.

It is a way of thinking, not a piece of software. It works with any tool, including a spreadsheet and a folder of bookmarks. Its guiding line is “Track who, not what.”

This manifesto makes the case in five theses. A [short version of the theses](https://picasi.com/en/manifest/) is published on picasi.com. This is the long version: each thesis with an explanation, the evidence behind it, and the strongest objection that can be raised against it.

## 1. Market knowledge is crucial, and market participants say a lot in public about their next steps

Anyone deciding on products, pricing or sales needs a picture of what competitors, customers, suppliers and industry associations are planning. Most of it is no secret. The people involved say it themselves, on their own channels.

Some of these publications are required by law. Under [Article 17(1) of the EU Market Abuse Regulation](https://publications.europa.eu/resource/cellar/b5b43c29-3d97-11f1-814f-01aa75ed71a1.0007.02/DOC_1), an issuer whose financial instruments trade on an EU venue “shall inform the public as soon as possible of inside information which directly concerns that issuer.” The same article requires the issuer to post that information on its own website and keep it there for at least five years. The company is the publisher of record for its own news.

Much of the rest is voluntary. The LinkedIn post by a competitor’s CEO, an association’s newsletter, a product video on a supplier’s YouTube channel, a job ad. Writing in [Harvard Business Review](https://hbr.org/2026/08/what-you-can-learn-from-a-competitors-job-postings), Wei Shi of the Miami Herbert Business School calls the jobs a competitor is trying to fill “one of the clearest and most accessible signals of strategic intent,” and notes that many companies overlook it.

All of these publications have one thing in common: you know who they come from. They are primary sources in the literal sense, the publication of the sender itself.

The obvious objection is that self-presentation is self-interested. A press release shows what a company wants to show, not necessarily what it does. That is true, and nothing in this approach denies it. Source-first means knowing who is speaking. It does not mean believing everything. How to weigh what a source says is the subject of the fourth thesis. One limit applies from the start: only publicly published content is in scope. What a company does not say stays invisible, and this approach makes no attempt to change that.

## 2. Social listening and keyword monitoring were designed for brand and topic monitoring, not for market analysis

Keyword monitoring starts with search terms. You get every hit that contains a word, no matter who used it. [Google’s help page on creating an alert](https://support.google.com/websearch/answer/4815696?hl=en) puts the principle plainly: “You can get emails when new results for a topic show up in Google Search.” The settings it lists are frequency, the types of sites, language, region, the number of results and which accounts get the alert. None of them is a setting for the sender. In fairness, Google Search, on which alerts are based, supports [operators such as site:](https://support.google.com/websearch/answer/2466433?hl=en) that restrict results to a specific site. That pins down the sender partly, but only by domain. A CEO’s LinkedIn post does not live on the company’s website.

In practice, social listening analyzes what third parties say about a brand, a product or a topic on social networks and the web, often with sentiment analysis. The starting point is the conversation about someone, not the statement by someone. Research uses the term more broadly. In the International Journal of Listening, [Stewart and Arnold (2018)](https://digitalcommons.unf.edu/unf_faculty_publications/1553/) propose defining social listening as “an emerging type of listening,” a way of gaining interpersonal information and social intelligence through mediated channels. The tools sold under that name, by contrast, are mostly built around brands and topics.

Both tools answer a legitimate question: where does a topic come up, and how are people talking about us? Market analysis asks something else. What is this competitor planning, what is this customer saying, what is this association preparing?

Information retrieval research explains why keywords handle the second question badly. The standard textbook by [Manning, Raghavan and Schütze](https://nlp.stanford.edu/IR-book/html/htmledition/information-retrieval-system-evaluation-1.html) states that relevance is assessed relative to an information need, not a query. A document is relevant if it addresses the information need, “not because it just happens to contain all the words in the query.”

In market analysis, the information need nearly always includes a sender. “What is competitor X planning?” is a different need from “Where does the word X appear?” A keyword search can hardly express the sender. At bottom, it matches words. Use it for this job anyway and you run into the familiar trade-off between precision and recall. The same textbook shows that the two [trade off against one another](https://nlp.stanford.edu/IR-book/html/htmledition/evaluation-of-unranked-retrieval-sets-1.html): you can always get a recall of 1 by retrieving every document, at very low precision. A broad alert delivers noise. A narrow one misses what the competitor says in different words.

A list of sources does not solve this with better keywords. It makes the sender part of the query.

The other side deserves a fair hearing. For crisis communication and reputation management, social listening is the right tool, because there what third parties say is exactly the point, and that is the question the tool was built for. Discovery works in a similar way. If you do not yet know who matters in a market, a topic search is often how you find new vendors and voices. Keywords help you find, sources help you follow. The full comparison is in [Sources vs. keywords](https://source-first-intelligence.com/en/sources-vs-keywords/).

## 3. AI-generated content is flooding the web, and content has become a mass-produced commodity

Language models can produce text in very large volumes. What that means for the total amount of content can so far only be measured in slices, and the slices point the same way.

In a study published as a preprint, Liang and colleagues examined [950,965 scientific papers](https://arxiv.org/abs/2404.01268) published between January 2020 and February 2024 on arXiv, bioRxiv and in Nature portfolio journals. Their statistical method estimates the share of text substantially modified or produced by language models at the level of whole corpora, not individual papers. They found a steady increase, largest in computer science at up to 17.5 percent. In [Findings of ACL 2024](https://aclanthology.org/2024.findings-acl.103/), Thompson and colleagues show that low-quality, machine-translated content makes up a large fraction of all web content in lower-resource languages.

None of these studies measures the share of machine-generated text on the web as a whole. Plenty of figures for that circulate. A figure for the web as a whole would need a measurement that none of these studies provides, so this manifesto cites none.

The consequences go beyond volume. In [Nature](https://www.pure.ed.ac.uk/ws/portalfiles/portal/460496122/ShumailovEtalNature2024AIModelsCollapseWhen.pdf), Shumailov and colleagues showed that indiscriminately training models on model-generated content “causes irreversible defects in the resulting models, in which tails of the original content distribution disappear.” They add that “it is unclear how content generated by LLMs can be tracked at scale.” The question of where a text comes from is open even for the people who build these models.

Lawmakers are responding. Under [Article 50(2) of the EU AI Act](https://publications.europa.eu/resource/cellar/b1730fb2-8f1c-11f1-9262-01aa75ed71a1.0001.03/DOC_1), providers of AI systems that generate synthetic audio, image, video or text content must mark the outputs in a machine-readable format. The obligation has applied since 2 August 2026. Under Article 111(4), systems placed on the market before that date have until 2 December 2026.

For market monitoring, this means more content does not bring more insight. It brings more filtering and more checking. Search by keyword and you get the growing volume delivered unfiltered.

One could object that competitors now have language models draft their press releases and posts too. Does the source still help? It does, for a reason the AI Act itself recognizes. When a deployer publishes AI-generated text on matters of public interest, Article 50(4) relieves it of the duty to disclose that the text was artificially generated if the text has undergone human review or editorial control and a person holds editorial responsibility for it. The provider’s machine-readable marking under Article 50(2) still applies. The law ties disclosure to whether someone stands behind the text. For market analysis, whether a model helped write a text is secondary. What counts is whether an identifiable sender stands behind it. A competitor’s press release remains the competitor’s statement, whoever drafted it. An anonymous text that happens to contain a keyword is nobody’s statement.

## 4. That’s why Source-First Intelligence starts with WHO, not WHAT: listen to the sources you can trust

The question is not “Who mentioned this keyword?” but “What is this competitor saying?” Asking it that way means deciding on the sources first and then reading their content.

Judging the sender and the message separately is not an invention of competitor monitoring. Military intelligence doctrine works this way. The UK Ministry of Defence doctrine [JDP 2-00](https://assets.publishing.service.gov.uk/media/653a4b0780884d0013f71bb0/JDP_2_00_Ed_4_web.pdf) (4th edition, 2023) describes the evaluation of incoming information as two judgments: the reliability of the source and the credibility of the information. “Reliability and credibility must be evaluated separately,” the doctrine says. In the NATO grading system, reliability runs from A to F and credibility from 1 to 6. A usually reliable source can supply inaccurate information, and highly credible information occasionally comes from a usually unreliable source. Reliability judgments depend largely on the source’s reporting history over time, and the doctrine explicitly includes organizational sources, not only individuals.

Content without an identifiable sender allows only half an evaluation. You can check whether a statement is plausible and whether others confirm it. You cannot check how reliable the sender has been, because there is none. And reliability only emerges from a history, which means following a source over time. That is exactly what Source-First Intelligence does.

“Trust” in this thesis therefore does not mean credulity. It means a source whose past statements you know and against which you can measure new ones. Suppose a competitor announces a product line for the third time and it again fails to appear. Then that competitor is a weak source for product announcements. Its job ads for a new plant may still deserve to be taken very seriously. You can only make that distinction if you know who is speaking.

A fair worry is that a fixed list of sources creates blind spots. What happens outside the list goes unseen. The risk is real. The doctrine requires gradings to be reviewed regularly. For a source list, practice adds a second step: adding new sources and dropping others. How to select, assess and maintain sources is covered in [Source-First in practice](https://source-first-intelligence.com/en/source-first-in-practice/).

## 5. The choice: track what just anyone says about your market, or what trusted market participants say themselves

The first four theses lead to a decision. It is less about tools than about where monitoring starts.

| | Starts with the keyword | Starts with the source |
|---|---|---|
| Guiding question | Where does the topic come up? | What is this market participant saying? |
| Entry point | Search terms, hashtags, brand names | A list of companies, people, associations and their channels |
| Result | Mentions from any source, including duplicates and spam | Content from known senders |
| Fits | Brand perception, crises, discovering new players | Competitor monitoring, market analysis, preparing decisions |

Suppose a maker of industrial pumps sets up an alert on the name of its main competitor. It receives hits about namesakes in other industries, dealer listings and job ads on recruitment sites. What it does not receive is the LinkedIn post in which the competitor’s CEO announces a move into maintenance contracts, because she writes “we,” not the company name. Had it followed the channels of that competitor and its management, it would have read the post the same day.

The choice is not either-or for all time. Often it makes sense to do both: social listening for your own brand, a source list for the market. What matters is not to confuse the two questions. Use brand-perception tools for market analysis and you get answers to a question you did not ask.

There is a second side to this thesis that is easy to overlook. Secondary sources have their place: reporting puts things in context, connects them and sometimes uncovers what companies would rather not say. The choice is not directed against them. It is directed against the habit of taking talk about a market participant for that participant’s own statement.

## What follows from this

The five theses lead to four subject areas that this site covers in depth:

- [Sources vs. keywords](https://source-first-intelligence.com/en/sources-vs-keywords/) sets Source-First Intelligence apart from keyword monitoring, media monitoring, social listening and alerts, and takes on the usual objections.
- [Source-First in practice](https://source-first-intelligence.com/en/source-first-in-practice/) describes how to select sources, find their channels, analyze their content and derive signals from it.
- [Source-First in context](https://source-first-intelligence.com/en/source-first-in-context/) places the approach alongside competitive intelligence, market intelligence, OSINT (open source intelligence, meaning insight from publicly available information) and media monitoring.
- [Source-First cases](https://source-first-intelligence.com/en/source-first-cases/) shows situations in which following sources makes the difference.

## Why the question of the sender will stay

Tools change. The question “Who is saying this?” stays, and it matters more the cheaper text becomes. A market is made of actors who decide, invest, hire and announce. If you want to understand what they are planning, listen to them. Not to everyone who talks about them.

Track who, not what.

## References

1. European Union: Regulation (EU) No 596/2014 on market abuse (Market Abuse Regulation), Article 17, consolidated version of 5 June 2026. [https://publications.europa.eu/resource/cellar/b5b43c29-3d97-11f1-814f-01aa75ed71a1.0007.02/DOC_1](https://publications.europa.eu/resource/cellar/b5b43c29-3d97-11f1-814f-01aa75ed71a1.0007.02/DOC_1) (accessed 23 September 2026)
2. Wei Shi: What You Can Learn from a Competitor’s Job Postings. Harvard Business Review, 10 August 2026. [https://hbr.org/2026/08/what-you-can-learn-from-a-competitors-job-postings](https://hbr.org/2026/08/what-you-can-learn-from-a-competitors-job-postings) (accessed 23 September 2026)
3. Google: Create an alert. Google Search Help. [https://support.google.com/websearch/answer/4815696?hl=en](https://support.google.com/websearch/answer/4815696?hl=en) (accessed 23 September 2026)
4. Margaret C. Stewart, Christa L. Arnold: Defining Social Listening: Recognizing an Emerging Dimension of Listening. International Journal of Listening 32 (2), pp. 85–100, 2018. DOI 10.1080/10904018.2017.1330656. Abstract in the University of North Florida repository: [https://digitalcommons.unf.edu/unf_faculty_publications/1553/](https://digitalcommons.unf.edu/unf_faculty_publications/1553/) (accessed 23 September 2026)
5. Christopher D. Manning, Prabhakar Raghavan, Hinrich Schütze: Introduction to Information Retrieval, section “Information retrieval system evaluation”. Cambridge University Press, 2008. [https://nlp.stanford.edu/IR-book/html/htmledition/information-retrieval-system-evaluation-1.html](https://nlp.stanford.edu/IR-book/html/htmledition/information-retrieval-system-evaluation-1.html) (accessed 23 September 2026)
6. Christopher D. Manning, Prabhakar Raghavan, Hinrich Schütze: Introduction to Information Retrieval, section “Evaluation of unranked retrieval sets”. Cambridge University Press, 2008. [https://nlp.stanford.edu/IR-book/html/htmledition/evaluation-of-unranked-retrieval-sets-1.html](https://nlp.stanford.edu/IR-book/html/htmledition/evaluation-of-unranked-retrieval-sets-1.html) (accessed 23 September 2026)
7. Weixin Liang et al.: Mapping the Increasing Use of LLMs in Scientific Papers. Preprint, arXiv 2404.01268, 1 April 2024. [https://arxiv.org/abs/2404.01268](https://arxiv.org/abs/2404.01268) (accessed 23 September 2026)
8. Brian Thompson, Mehak Dhaliwal, Peter Frisch, Tobias Domhan, Marcello Federico: A Shocking Amount of the Web is Machine Translated: Insights from Multi-Way Parallelism. Findings of the Association for Computational Linguistics: ACL 2024, pp. 1763–1775, August 2024. [https://aclanthology.org/2024.findings-acl.103/](https://aclanthology.org/2024.findings-acl.103/) (accessed 23 September 2026)
9. Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Nicolas Papernot, Ross Anderson, Yarin Gal: AI models collapse when trained on recursively generated data. Nature 631, pp. 755–759, 24 July 2024. DOI 10.1038/s41586-024-07566-y, version of record in the University of Edinburgh repository. [https://www.pure.ed.ac.uk/ws/portalfiles/portal/460496122/ShumailovEtalNature2024AIModelsCollapseWhen.pdf](https://www.pure.ed.ac.uk/ws/portalfiles/portal/460496122/ShumailovEtalNature2024AIModelsCollapseWhen.pdf) (accessed 23 September 2026)
10. European Union: Regulation (EU) 2024/1689 (Artificial Intelligence Act), Article 50 and Article 111(4), consolidated version of 27 July 2026. [https://publications.europa.eu/resource/cellar/b1730fb2-8f1c-11f1-9262-01aa75ed71a1.0001.03/DOC_1](https://publications.europa.eu/resource/cellar/b1730fb2-8f1c-11f1-9262-01aa75ed71a1.0001.03/DOC_1) (accessed 23 September 2026)
11. UK Ministry of Defence: Joint Doctrine Publication 2-00, Intelligence, Counter-intelligence and Security Support to Joint Operations, 4th edition, August 2023, paragraphs 3.39 and 3.40. [https://assets.publishing.service.gov.uk/media/653a4b0780884d0013f71bb0/JDP_2_00_Ed_4_web.pdf](https://assets.publishing.service.gov.uk/media/653a4b0780884d0013f71bb0/JDP_2_00_Ed_4_web.pdf) (accessed 23 September 2026)
12. Google: Refine Google searches. Google Search Help. [https://support.google.com/websearch/answer/2466433?hl=en](https://support.google.com/websearch/answer/2466433?hl=en) (accessed 23 September 2026)
