Private data profiler
Know what's in your CSV before you trust it.
Drop in a file and see column types, distributions, missing values, duplicates and data-quality problems. It runs in your browser, so the file never leaves your machine.
No sign-up. The sample is 50,000 generated rows.
- WARNDuplicate rows68 rows (0.14%) are exact copies of an earlier row.
- WARNunit_price / Values that aren't numbers37 values can't be read as numbers, though the column looks numeric.
- NOTEregion / Possible placeholder"N/A" appears 300 times and may stand for a missing value.
- NOTErating / Missing values35% of values are missing (17,710 rows).
- MEDIAN
- 44.29
- MEAN
- 72.86
- MAX
- 823.74
- 0
- bytes uploaded
- 10 MB
- streaming chunks
- 27
- engine unit tests
- MIT
- open source license
01 / How it works
Three steps, and none of them is an upload.
- 01
Drop in a CSV
Or open the built-in sample. Column types are detected from the first rows straight away.
- 02
Read the profile
Findings come first, then per-column statistics, distributions and the most common values.
- 03
Click to explore
Filter to any value and the whole profile re-runs on just those rows. Export a report when you're done.
02 / What it checks for
The problems that quietly skew a chart or break an import.
- Duplicate rows
- Exact copies of an earlier row, counted across the whole file.
- Values that aren't numbers
- Text hiding in a column that otherwise looks numeric, like a stray "N/A" price.
- Missing data
- Empty cells per column, flagged when a column is mostly empty.
- Placeholders
- Values like N/A, null or - that usually stand in for something missing.
- Outliers
- How much of a column falls outside the usual range, using the 1.5×IQR rule.
- Constant and ID columns
- Columns with a single value, and columns where every value is unique.
03 / Private by design
The site has no upload endpoint at all.
Your file is read and analyzed inside your browser tab. There is no server that receives files, so sensitive exports like HR sheets, finance data and customer lists stay on your machine. Close the tab and it's gone.
04 / Under the hood
Rust, compiled to WebAssembly, streaming the file in 10 MB chunks.
- Web Worker
- The engine runs off the main thread, so the page stays responsive on large files.
- Chunk-safe parsing
- Rows split across chunk boundaries, newlines inside quoted fields and a last row with no trailing newline are all handled, and covered by unit tests.
- One-pass statistics
- Mean and variance use Welford's algorithm. Medians, quartiles and histograms come from a 2,000-value reservoir sample.
- Bounded memory
- Top values are capped at 1,000 distinct entries, and duplicates are tracked with 64-bit row fingerprints instead of the rows themselves.
05 / Good to know
What it doesn't do yet, so you aren't surprised.
- CSV only for now: comma-separated, with a header row. No Excel, JSON or Parquet yet.
- On large files, medians, quartiles, outlier rates and histograms are estimates from a sample. Counts, minimum, maximum and mean are exact.
- Above 1,000 distinct values a column is marked high-cardinality and its top-value counts are a lower bound.
- Dates are recognized in YYYY-MM-DD format only.
See it on 50,000 messy rows.
The sample has duplicates, bad values and missing data planted in it, so you can watch the profiler catch them.