What data about you is already available and what you can do about it
A story from February 2026 about an open-source project explicitly conducting surveillance tech is recirculating on social media, as the dark history and work of data synthesis company Palantir becomes more and more controversial and gains more and more attention.
OpenPlanter was released around February 2026 by a pseudonymous developer going by “Shin Megami Boson” as an “open source option” being positioned by some of the discourse as a way to connect the dots on right wing dark money and actions. It’s a recursive-language-model investigation agent that breaks large investigative objectives into sub-tasks via delegated sub-agents (default depth of four), resolves entities across datasets with no common identifiers, and builds evidence chains for its findings , pitched, partly characterized by MarkTechPost as “helping you keep tabs on your government since it’s keeping tabs on you.”
The initial coverage wave, led by MarkTechPost’s “community edition of Palantir” framing, set the discourse template that’s persisted since: a “democratization of surveillance” narrative in which smaller organizations and individuals gain investigative capabilities previously restricted to well-funded institutions, usually paired with a hedged acknowledgment that this raises governance questions around privacy and the ethical deployment of micro-surveillance. Notably, even MarkTechPost attached a disclaimer that it does not endorse the project, and a minority thread in the commentary pushed past the mini-Palantir narrative to architecture questions.
One analysis argued the edge-versus-cloud deployment choice will determine whether it minimizes data exposure or replicates the centralized surveillance model it claims to disrupt. Since then the project has cycled through periodic viral resurrections on X, with the dual-use and reliability critiques living mostly in the replies rather than the quote-tweets.
“Open-sourced Palantir” is not really accurate, a category error. Palantir’s moat was never just software, it’s the contracts, the privileged data access, the forward-deployed engineers embedded inside agencies, the institutional integration. You can clone the interface pattern (entity resolution + knowledge graph + provenance) but not the position. Calling it “Palantir for free” flatters both projects: inflates OpenPlanter and launders Palantir into “just a product” rather than a political arrangement.
The symmetry framing is doing a lot of work. “They investigated us, now we investigate them” assumes the tool is directional. It isn’t. An LLM-driven entity-resolution engine pointed at “public documents” works exactly as well for doxxing an activist, mapping a union drive, or building a harassment dossier as it does for accountability journalism. Diffusing capability isn’t the same as redistributing power: institutions keep their proprietary data, legal access, and compute, while individuals gain exposure.
The reliability problem is buried. LLM entity resolution across messy datasets is confident conspiracy-graph generation with a false-positive problem: it can merge people who share names and surface “non-obvious connections” that are statistical noise. “You can inspect the original sources” is carrying the entire epistemics of the pitch, and it assumes users will. About real, named people, that’s a defamation machine as much as an investigation tool.
What Palantir actually sells. The software layer (ontology management, entity resolution, link analysis, provenance tracking) has been semi-commodity for a long time. IBM’s i2 Analyst’s Notebook did link charts in the 90s; Maltego has done OSINT graphing for two decades; graph databases are open source. What Palantir monetizes is everything around the software: forward-deployed engineers who embed inside the client for months and hand-build the data integrations, the security clearances, the ATO certifications that let it run on classified networks, and above all the contracts themselves, often sole-source procurement relationships with ICE, DoD, NHS, police forces. The “product” is an institutional arrangement with a UI on top. This matters beyond pedantry: if you believe Palantir is software, you conclude the fix is competing software. If you understand it as a procurement and data-access arrangement, you conclude the fix is procurement reform, data governance, and contract transparency, which is exactly the terrain incumbents prefer critics stay off. The “open-source Palantir” frame concedes their preferred definition of themselves.
Why the symmetry doesn’t hold. The tweet is running a sousveillance argument (Steve Mann, David Brin’s “transparent society”) that watching the watchers rebalances power. The problem is that surveillance capacity has three components: data access, analytic capability, and the ability to act on findings. OpenPlanter diffuses only the middle one. The state still has subpoena power, data-broker purchasing, classified feeds, and, critically, enforcement capacity. A citizen who maps a corruption network still needs a prosecutor, a newsroom, or a court to make it matter; an agency that maps a citizen can act directly. Meanwhile the diffusion cuts downward more sharply than upward: investigative journalists already had OSINT skills and tooling, so their marginal gain is modest. The person who wants to build a dossier on an ex, an activist, or a local official’s daughter gains capability they never had. There’s also a doctrinal point buried here: aggregation is itself the harm. US law partially recognized this in Supreme Court decision DOJ v. Reporters Committee. “Practical obscurity” means scattered public records are meaningfully private until compiled. A tool whose entire function is compilation dissolves that protection for everyone, in both directions. “It’s all public documents” is the same defense the data brokers use.
The failures get worse under recursion. Classical entity resolution has known, measurable error rates and conservative matching thresholds. LLM-based resolution can often replace thresholds with plausibility, and language models are optimized to produce coherent narrative, which is the failure mode you definitely don’t want in an investigation tool. A system that “surfaces non-obvious connections” cannot distinguish between a hidden network and a coincidence it narrated well. The recursive architecture can compound this: once a false merge enters the shared context or a written artifact, later subtasks may inherit it as an established premise, so a bad merge at depth one can become settled fact by depth four. The “evidence chains behind every finding” feature is real but functions as provenance theater in practice: automation bias is well documented, and a confident graph with citations attached gets checked less, not more, because the citations signal rigor. And a graph visualization is rhetorically loaded in a way prose isn’t: two nodes with a line between them looks like an accusation. The people harmed by false merges will be real and named, and the liability often sits with whoever publishes, not with the repo.
Why the genre matters. The February launch resurfacing in July as “just happened” isn’t trivial; the viral-AI-tool format (mystery founder, staccato line breaks, “the idea is huge”) is how tools get pre-framed before anyone critical arrives. By the time governance people engage, “democratizing Palantir” is the established vocabulary and skepticism reads as defending Palantir. That’s the same move as “democratizing AI” generally: a distribution claim smuggled in as a liberation claim.
The accurate version is still a big story. Intelligence-analysis workflow, the actual craft of fusion centers and corporate intelligence shops, is becoming something that runs autonomously on a laptop. That’s significant, and it cuts into procurement debates directly: governments justify vendor lock-in partly on the premise that these capabilities are scarce and specialized. If they’re commodity, the sole-source rationale weakens. But the same commodification means every capability argument made about state surveillance now applies to a Docker container with shell execution. The interesting piece isn’t “the public gets Palantir.” It’s that the barrier between institutional and individual intelligence capacity is collapsing, and nobody’s governance framework, “theirs” or “ours”, was written for that.
What Data is Available to Whom
It varies by jurisdiction, but here’s a good guideline: Availability depends on who is asking, so the important question is who can get it, since “available” means different things for a stranger with an agent, a data broker client, and a government.
Tier 1: Available to anyone with a laptop. Government filings (property, corporate, court, donations, licenses, liens), breach dumps (passwords, emails, phones, addresses from two decades of leaks), scraped social history (every public post, like, follow, and deleted-but-archived tweet), and people-search aggregates built from all of the above. This is the OpenPlanter tier. Assume full availability.
Tier 2: Available to anyone with a modest budget. The ad-tech layer. Bidstream data leaks your location, device ID, and app usage to hundreds of companies every time an ad auction runs, and brokers resell it with minimal vetting. Mobile advertising IDs link your physical movements to your interest profile. This is how journalists deanonymized a priest via Grindr data, and how agencies including ICE, the FBI, and DHS have bought location data without warrants. So yes, ad data linked to government records is a real, operating pipeline, purchasable rather than public.
Tier 3: Held by platforms, reachable by legal process, leaks, or policy change. Browsing history (your ISP and Google hold it; strangers cannot buy it directly, though the ad-tech shadow of it approximates it), search history, and AI prompt history. Prompt history deserves special attention: it is the most confessional dataset that has ever existed, it is held by a handful of companies, it is subpoenable, and it has already surfaced in criminal cases. The OpenAI–NYT litigation forced retention of chats users believed deleted. Assume prompts are discoverable records, and assume the retention promises attached to them are contingent.
Tier 4: Algorithmic inferences about you. Credit scores, insurance risk scores, tenant-screening scores, fraud scores, ad-interest categories (“expectant parent,” “diabetic interest”), and increasingly LLM-derived profiles. These are the least visible and least contestable tier. They are mostly proprietary, so a stranger cannot pull them, but they shape prices, housing, and eligibility decisions about you daily, and inference regulation lags collection regulation badly. The near-term risk is that agentic tools reconstruct approximations of these scores from Tier 1 and 2 data, which puts institutional-grade profiling in the stranger tier too.
Your phone number: arguably the single worst identifier you have, because it functions as a universal join key. It sits in Tier 1 through breach dumps and people-search sites, gets sold in Tier 2 broker data, and is the field entity-resolution tools love most, since unlike names it’s globally unique. One number can stitch together records that share nothing else
What it links: every account using it for 2FA or recovery, your messaging apps (WhatsApp, Signal, Telegram all key on it), and everyone who uploaded their contacts to any app, meaning your number-to-name mapping was leaked by your friends’ address books years ago regardless of your own hygiene. The 2021 Facebook scrape alone tied 533 million numbers to identities, permanently. Reverse lookup from number to name, address, and relatives is a commodity service. It’s also carried in ad-tech and loyalty-program data, so it bridges your commercial profile to your legal identity.
It’s an attack surface too, not just an identifier. SIM-swapping turns your number into account takeover, which is why SMS 2FA is the weakest form.
What you can do: treat your real number like an SSN/SIN. Use a VoIP or secondary number (Google Voice, MySudo) for signups, commerce, and anything public-facing, and reserve the carrier number for banks and people you know. Switch 2FA to authenticator apps or passkeys so the number stops being a recovery key. Set a carrier PIN against SIM swaps. Enable Signal’s phone-number-privacy setting so contacts can’t harvest it. The catch is the one you’d expect: the number is already out there from past decades, so this limits future linkage rather than erasing past linkage.
The compounding effect matters more than any tier alone. Tier 1 gives a skeleton, Tier 2 adds movement and behavior, Tier 3 waits in legal reserve, Tier 4 converts it all into decisions. What changed in 2026 is that the labor cost of assembling tiers into a single coherent file dropped to a prompt.
Assumptions to make about your data
- Assume everything filed with a government is aggregatable. Property deeds, corporate registrations, court records, campaign donations, professional licenses, voter rolls in many jurisdictions. Each record was designed to be obscure in isolation. Tools like this exist to end that isolation, and the practical obscurity that once protected scattered records is gone.
- Assume your name links your records. Entity resolution works even without common ID numbers. A donation under your name, a deed under your name and a spouse’s, an LLC where you’re listed as director: these now resolve to one node in a graph, along with the false matches that come with sharing a name.
- Assume data brokers already sold the rest. Location history, purchase patterns, household composition. Anything a broker sells to institutions can eventually be bought, scraped, or leaked into a dataset someone feeds an agent.
- Assume breach data circulates permanently. Old passwords, addresses, phone numbers from a decade of leaks remain in dumps that ingestion tools handle as easily as clean CSVs.
- Assume the audience is anyone, for any reason. The person running the query can be a journalist, an ex-partner, an employer, a stalker, or an obsessive stranger. Capability diffusion means threat modeling can no longer start from “who would bother.”
- Assume inferences will be wrong and confident. You can be misidentified, merged with a stranger, or placed in a “non-obvious connection” that is statistical noise rendered as a chart. The graph about you may be defamatory before it is accurate.
What you can do to manage data risk and exposure
- Separate your names. Use an LLC or trust for property where lawful, a PO box or registered agent for business filings, and distinct emails and handles across contexts. Linkability is the attack surface; break the links you control.
- Send removal requests to major brokers, or pay a deletion service to do it on a schedule. Removal is a treadmill, and a treadmill still beats standing still.
- Audit what a stranger sees. Search your own name, address, and phone number periodically, including in breach-check services, so you learn what a graph of you contains before someone else does.
- Restrict what only exists because you posted it. Public records are hard to claw back; social media, resumes, and forum history are the discretionary layer of your file.
- Know your jurisdiction’s rights. GDPR erasure, CCPA deletion, and Canada’s PIPEDA complaints all move slowly and unevenly, and they remain the only mechanisms with legal force behind them.
- Treat the structural fixes as political. Individual hygiene manages exposure at the margin. The record-keeping regimes, broker markets, and procurement arrangements that create the exposure can only be changed collectively, which is the limit of any list like this one.

