See the accounts shaping your feed before they shape your views.
An AI copilot for public discourse · by SimPPL, a U.S. 501(c)(3) nonprofit
of Americans trust the mass media, the lowest in Gallup's five-decade trend (Sept 2025).
worldwide worry about what is real and what is fake in online news (Reuters, 2025).
of Americans get news via social media and video, ahead of TV for the first time (Reuters, 2025).
AMERICANS WITH A GREAT DEAL OR FAIR AMOUNT OF TRUST IN MASS MEDIA · GALLUP
Jobs, health, elections: real, documented cases, sourced on every card.
On Twitter, true stories took about six times as long as false ones to reach 1,500 people (MIT study, 2018). By the time a claim is checked, the campaign that planted it has moved on.
We rely on journalists and fact-checkers to catch this, but they've lost their view: CrowdTangle shut down in 2024, and paid access now costs thousands a month. And each platform moderates only its own feed, so a campaign across X, YouTube, Telegram, and WhatsApp shows each watcher a quarter of the picture.
OnPassive, an "AI-powered" tech scheme, recruited across social platforms at once. The SEC charged it with fraudulently raising over $108 million from more than 800,000 investors. Each platform saw only its own slice.
We help collect public social data from sources within platform terms of service, including platform-granted APIs, across X, YouTube, Reddit, Bluesky, Instagram, Facebook, TikTok, LinkedIn, Telegram, and 4chan, plus 20,000+ newspapers and Wikidata.
Posts cluster into themes and sub-themes. Each theme surfaces the accounts moving it, with per-actor dossiers and influence maps.
Deep research agents answer your questions across the corpus and cite every post they draw on. You keep the judgment.
Pick a topic or a set of accounts and a date window, and Arbiter compiles a month of cross-platform conversations in 20 to 40 minutes. We've built reports on over 30 issues this way, from job scams and human trafficking to the manosphere influencers reaching younger audiences.
Every study keeps its sources attached: each number traces back to the posts behind it.


Posts cluster into themes, themes into sub-themes, and each cluster names the accounts driving it. The graph turns a feed into a map: who amplifies whom, in which language, on which platform.
Reports built this way sit behind both platform outcomes on the track-record slide: X's internal investigation and Meta's Bangladesh takedown.
We put this to a live study of the manufactured attention around Ibrahim Traoré, the officer who seized power in Burkina Faso. The agent read all 191 posts and ranked the accounts driving the conversation.
You interrogate a corpus too large to read in plain language and get an answer you can act on. The engagement traces to a handful of accounts spiking in single bursts rather than sustained volume, the shape of a coordinated push, not a real audience.
Every row ties back to the posts behind it. A question like this costs 5 credits and takes under a minute.
Engagement is driven by a small set of accounts, mostly via high-performing posts that don't include outbound links: native video and text rather than link-sharing.
| Account | Posts | Interactions | Outbound links |
|---|---|---|---|
| zoomafrika1 | 1 | 56,860 | none detected |
| SahelAlerte | 1 | 29,721 | none detected |
| WelcomeTheGulag | 1 | 23,175 | none detected |
| kelevitch | 1 | 22,426 | none detected |
| WithoutHistory | 1 | 10,499 | none detected |
| kwadwosheldon | 1 | 8,725 | youtube.com |
Search the posts on every platform, aggregate the ones that move together, then investigate the accounts behind them.
The language that carries a story is coded, multilingual, and shifts week to week. The hard part is retrieving what someone intended even when the obvious keywords would miss it.
A feed of millions of posts hides its own structure. Turning it into a navigable map of narratives and the accounts moving each one, resolved to real entities, is more than clustering.
Ask any question of the corpus and an agent writes and runs the analysis, returning a cited answer. Surfacing custom insight on demand is the value; grading each answer against baselines is what makes it safe to rely on.
Most listening tools answer with RAG over whatever their feeds already collected. Arbiter plans the collection itself and surfaces conversations you did not know to ask for. The next slides show the method; the appendix lists every model.
Your question expands into a typed plan: the actors, places, and domains it implies, from a daily-refreshed knowledge graph.
Keyword and meaning search, fused, then reranked to what actually answers. One threshold drops 60 to 80% of the noise.
No one labels their own post skeptical. Search the word skeptic and you find almost nothing. The pushback itself, "the models always run hot," "ice has advanced before," shares none of your words, and Arbiter surfaces it by meaning.
A keyword tool returns matches. In a separate April 2026 window we asked one question about Ibrahim Traoré, collected 22,719 posts, and kept the 19,532 relevant ones. Arbiter aggregates them into a map you can walk, concept to sub-concept to post, including conversations the query never mentioned. click any row below
The YouTube tab of this study holds 51 concepts, 41 sub-concepts, and 3,817 posts from 1,115 users; this slide walks one path. The rest open in the live study.
Say a search comes back with 100,000 posts. Point an LLM at them naively and the bill explodes while the labels collapse into duplicates: 28 of 53 clusters in one corpus came back with the same label.
We group the posts by meaning first, in a geometric space, before any LLM reads them. It is token-efficient where frontier-model approaches are not, a first-of-its-kind method we were invited to present at Stanford.
The LLM then labels whole clusters, so cost scales with a few dozen clusters, not a hundred thousand posts. Any label that lands too close to an existing one is thrown out.
The LLM provides the vocabulary. The geometry enforces distinctness.
The climate query on arbiter.simppl.org/features: naive labeling collapses to one theme; ours keeps them distinct. (Bars: average pairwise label similarity, shorter is better.)
The same embedding space powers entity extraction: nine entity types, resolved to Wikidata, name variants merged automatically. Each actor gets sentiment, emotion, and a clickbait score on a five-point scale.
Arbiter sits as an analytical layer over social platforms and online sources. These signals add up to tactics: manufactured consensus, coordinated amplification, borrowed legitimacy, undisclosed promotion. We test the classifiers in non-Latin scripts too, from Hindi to Cyrillic, so manipulation in those languages isn't missed.
Ask a question and the agent runs the investigation. It picks from 30+ analysis functions, then writes and runs code against the corpus in a sandbox.
Every answer arrives as an artifact you can verify: the code it ran and the source posts it cites.
Every answer is also graded automatically, and the error patterns feed continuous improvement, so each fix reaches every future query.
You keep the judgment. The agent does the digging.
Every line links back to the source posts. From the "Narratives regarding Climate change" study on arbiter.simppl.org.
Collecting posts is the easy part. The hard part, and what we build, is turning them into evidence you can trust. Five problems make it hard:
Coordinated pages promote Ibrahim Traoré with sentiment reading 100% positive, 100% joy. Organic discourse rarely reads this way.
One account lands 88.8K interactions from a single post (views, comments, and likes combined). Its clickbait score is a clean 1 of 5. The campaign is built to look like ordinary fan praise, and to a casual reader it does. Arbiter flags it anyway: organic audiences disagree with each other, and this one never does.
Fake factories, fake railways, fake tributes, drawing millions of views before any platform labels them.

Uniformly green panels are the anomaly that starts the investigation.
One year on (May 26 to Jun 2, 2026), a rebuilt network of 30 channels cut from one naming template works the same beat, now with tip jars, memberships, and merch links attached. One of the 2025 actors is back under the same name.

"This video is no longer available because the YouTube account associated with this video has been terminated."
Two studies built with Migrasia read 2,251 recruitment posts from 1,444 accounts. Most scam posts work hard to look legal: 489 Indonesian posts open with words like "resmi" (official) and "no upfront fee", and 183 Filipino posts print a recruitment licence number.
The contact details give the networks away: one shortened link posted by 21 regional job pages under different names; six profiles that all route to a single WhatsApp number; in the Philippines, 34 shared phone numbers and emails tie the pages into 14 networks.
One agency kept posting vacancies under licence DMW-011-LB-051822-R after that licence had expired, so the licence number itself was part of the sales pitch.
Recruitment tactics by post count, Indonesian corpus (1,492 posts); bars scaled to the leader.
A quarter of the Filipino corpus is the aftermath: airport interdictions, arrests for illegal recruitment, repatriations of trafficking victims.
Two months of high-engagement posts drew 127.3M interactions. The content is sold as self-improvement: "looksmaxxing" routines and testosterone regimens, with verified health information pushed aside.
Engagement runs on a star system: the top 3 accounts take 14.6% of all engagement, while the highest-volume cluster (1.8K posts) registers 0.0%. Culture-war framing carries 57% of interactions; belonging and dating discourse carries 56.8%.
The report maps outbound audience redirection: the links actively shared to move audiences off-platform, and the content driving them there.
Share of total interactions carried by each theme (themes overlap, so shares exceed 100% combined).
Inside the conversation, a few accounts like @DrewPavlou and @MikhailaFuller carry it, each turning one or two posts into millions of interactions. The recurring claims run from testosterone decline to dating advice, wrapped in culture-war and media-accountability framing; Netflix's documentary was one entry point.
A vendor printing error sent some voters the wrong party's primary ballot. Unable to identify who was affected, the state re-sent replacements to everyone mailed a ballot before May 14, covering 500,000+ requests. Originals were voided; every return envelope carries a unique identifier, so officials state there is no risk of duplicate voting.

The official explanation, cited in the study.
The stakes were legislative: the study's top post (1.32M interactions) demanded the Senate attach the SAVE America Act, which would require in-person documentary proof of citizenship to register and would replace automatic mail-ballot delivery with application-only mail voting. The false ballot claims circulated as the Senate weighed it; the act passed the House 218–213 and failed in the Senate.
We tested our own pipeline against hand-labeled data before trusting it. Two results changed how we build.
The agent writes and runs TypeScript in a sandbox against the corpus, picking from 30+ analysis functions, capped at eight steps. Every answer returns the code it ran and the posts it cites, and any run replays deterministically.
A tip, a beat, a policy window. We scope the corpus, platforms, and dates together.
Retrieval, themes, and actor dossiers land in a case-study library. Your analysts question it through the agent, with a citation on every answer.
The findings are yours to publish. Evidence stays attached, so editors and reviewers can audit every number.
One scoped corpus replaces manual sweeps across four platform search bars.
In Mongolia, NEST read a defamation bill's direction in the data 19 days before an eyewitness confirmed it.
Coordinated campaigns surface before they reach your abuse reports.
Our scored corpora and dossiers slot in as an analytical layer under your platform.
Model stages are commodity parts; the retrieval planning, the thresholds, and the grading harness are ours. When a better model appears, we swap the part and your workflow does not change.
100+ organizations already use Arbiter, on five continents.
Newsrooms, fact-checkers, UN trust-and-safety teams, and U.S. state election officials. YouTube grants us 1.2M API calls a day. Stanford picked Arbiter for this year's platform workshop; Meta, YouTube, and TikTok held that slot in past years.
Former postdoc in AI safety at MIT and Boston University. Board member at Integrity Institute. Built AI systems at Twitter, Adobe, and Slack. Google Research Innovator.
Led audience-analytics ML as Data Scientist III at Bombora. M.S. in Data Science at NYU. Fellow with the Center for AI and Digital Policy and at the University of Mannheim.
System design and infrastructure, NLP and deep research agents, platform backend and APIs.
14 peer-reviewed papers (NeurIPS, ICML, AAAI, ICWSM). Awards from Google, Mozilla, Wikimedia, Ford, and Omidyar. Bootstrapped as a nonprofit since 2021.
Collects public social data within platform terms of service, clusters themes, maps actors and networks, and answers questions with a citation for every claim.
Draws the conclusions. Journalists, researchers, and trust & safety teams keep the judgment; Arbiter keeps the evidence attached.
We fact-source: the system surfaces claims and groups related ones, so one verified fact-check scales to thousands of posts. We spent a year building this with Wikimedia.
The case files in this deck are public studies from 2025 and 2026 windows. A study scoped to your issue tracks your actors, in your languages, on your timeline.
Each activity is an annual tier you can lead on its own. Fund all three for a year: $1M. The full 24-month proposal is $1.8M. A scoped pilot on your own issue starts at $25K–$150K; anchoring a new country vertical runs $1.5M–$3M.
Grants subsidize free access; paid subscriptions with larger newsrooms and civil-society clients grow each year. first recurring revenue in 2025 at $15K a quarter. Since 2021 we have brought ~$450K to SimPPL from grants raised jointly with collaborators.
Request a demo: write to team@simppl.org
and we follow up within the week.
A pilot is one case study scoped to your issue: we set the date window together, run the investigation with you, and you keep the evidence and the sources. Our grants subsidize free access for independent journalists and fact-checkers, and paid subscriptions with larger newsrooms and civil society clients generate growing revenue each year to support our moonshot projects and team members sustainably. Free investigation reports land in your inbox via simppl.org/newsletter. The platform is live at arbiter.simppl.org.