PII Discovery
Finds personal data where it actually is, across databases, shares, mailboxes, drives, APIs and endpoints, and turns it into the Record of Processing, the DPIA, the erasure and the breach scope.
The demo runs on synthetic data. A read only visitor account is published here shortly; until then, request a demo and we open the console on your own source.




What it does
- Scans uploaded CSV, TSV, Excel, PDF, Word, text, Markdown, log, JSON and .eml files; PostgreSQL, MySQL and MariaDB, SQL Server, Oracle and SQLite, with Snowflake, Redshift, BigQuery and Synapse by adding their driver; MongoDB as nested paths; folders and network shares in place; Gmail, IMAP and Microsoft 365 mailboxes read only; any REST API returning JSON; OneDrives; a whole Microsoft 365 tenant in one resumable sweep; Windows endpoints.
- Ten detector groups: core identity, contact, government and tax IDs (Aadhaar with checksum, PAN, passport, voter ID, driving licence, vehicle, GSTIN), financial and payment (Luhn checked cards, UPI, IFSC, IBAN, accounts), employment (UAN), special category, health and clinical (ABHA, UHID, MRN), student and academic (APAAR, ABC, UDISE), banking and lending (CIF, CIBIL, demat), insurance (policy, claim, TPA). Names, addresses and organisations in free text by language model.
- A data map with type, category, sensitivity, count, a masked sample and where the data is held. Raw personal data is never stored; passwords and tokens are used for the scan and discarded.
- Who holds data: one row per person and per machine, with an erasure worklist as CSV.
- RoPA pre filled from discovery: purpose, lawful basis, fiduciary, processors, categories of data principals, sensitive categories, recipients, retention, cross border transfers, security measures; exported to Word and Excel.
- DPIA and PIA workflow with the inventory seeded from discovery, a risk register, mitigations, sign off and privacy threat modelling on LINDDUN categories; a data flow map with cross border transfers highlighted.
- Subject search and estate wide erasure and correction: files the platform holds are quarantined reversibly with a restore during the grace period; sources it can only read become tracked manual actions verified by a fresh scan.
- Scheduled rescans with drift alerts, exposure findings (open shares, PII in logs and exports, data past retention), redaction and pseudonymisation on export, an access request pack per person, SIEM and ticketing hooks.
- Owner, analyst and viewer roles, MFA, SSO, and an audit trail of every scan, export and change.
How it maps to the law
| Obligation | Where | What the product does | Status |
|---|---|---|---|
| Know what personal data you process | Section 8(1) | Discovery across the estate with validated Indian identifiers. | Built |
| Erasure on withdrawal or when the purpose is served | Section 8(7) | Estate wide erasure, reversible quarantine, rescan verification. | Built |
| Rights of access and correction | Sections 11, 12 | Subject search and the correction workflow, opened from the consent platform. | Built |
| Reasonable security safeguards | Section 8(5), Rule 6 | Masked samples only, credentials never stored, exposure findings. | In build, October 2026, exposure findings |
| Record of processing for a Significant Data Fiduciary | Section 10, Rule 13 | RoPA pre filled from discovery, exported for audit. | Built |
| Data protection impact assessment | Section 10(2), Rule 13 | DPIA workflow with threat modelling and sign off. | In build, October 2026 |
| Cross border transfer restrictions | Section 16, Rule 15 | Location held per source; transfers flagged in the RoPA and the flow map. | In build, October 2026, flow map |
| Breach intimation | Section 8(6), Rule 7 | See the Breach Register. | Built |
Honest limits
- Discovery samples the first rows of a structured source by default, which is fast and reliable for classifying columns but can miss rare values; use the full scan option to read every row. Documents are always read in full.
- Name and address detection depends on the language model, which the Docker build installs. Outlook .msg files are exported to .eml first.
- It is an aid to data mapping, not a guarantee of completeness, and its output is a finding to be reviewed.
Deployment
FastAPI, PostgreSQL, React, spaCy, Microsoft Graph. Self hosted in your India region cloud, next to the data, or hosted by us. Read only, least privilege connections throughout.
Scan one source in the demo
Upload a CSV or connect a read only database, open the data map, and generate the RoPA from it.