A CSV with a million rows: what to open it with
Excel's limit is 1,048,576 rows. A file with two million rows opens exactly half-way, and the warning is easy to miss — which is more dangerous than refusing to open at all.
Free for personal use. Files are read in parts, in your browser — nothing is uploaded. Windows version — 3.5 MB, no installation: details. Organizations: license.
First, do not lose data
The main danger of a big CSV is not that it is hard to open, but that it is easy to open partially without noticing. Before you calculate anything, make sure you see the whole file: compare the number of rows in the source with the number you see.
A quick way to get the exact row count without opening the file: on Linux and macOS, wc -l file.csv; in PowerShell, (Get-Content file.csv | Measure-Object -Line).Lines. Remember that the header counts too, and that a line break inside a quoted field inflates the result.
The options, laid out honestly
| Method | When it fits | What it costs |
|---|---|---|
| Power Query in Excel | You need to aggregate and build a pivot table | Data is loaded into a model and memory grows; browsing row by row is awkward |
| Python + pandas | A one-off job, and someone can write the code | The whole file goes into memory; a 5 GB file needs 15–20 GB of RAM |
| DuckDB | You want SQL over the file and are fine with a command line | An excellent choice, but a tool for people who write queries |
| A text editor (Notepad++, VS Code) | Peeking at the first lines | No columns, no filters, and it crawls at a gigabyte |
| A browser-based viewer | Look, filter, reconcile, export a subset | Not a replacement for an analytics tool |
If the task sounds like “look and find”, heavy machinery is not needed. If it sounds like “recalculate the year's figures”, it is — and that means DuckDB or a database.
Why a big CSV opens slowly in the first place
A CSV has no table of contents. To know where the millionth row starts, you have to read all the rows before it: a line break inside a quoted field does not end a record, so “just search for \n” does not work.
That leaves two approaches. The first is to read the whole file into memory and then work fast (pandas and Excel do this). The second is to walk through the file once, remember the offset where each row starts, and read the rows themselves on demand. The second path needs one sequential read and hardly any memory: for ten million rows the index takes about 80 MB, while the data itself stays on disk.
Tabulens takes the second path — so a gigabyte file opens in about the time the disk needs to read it, and scrolling does not depend on size.
What else breaks on big files
- Ragged rows. Among a million rows there will almost certainly be a dozen with an extra delimiter. A tool should report it, not silently shift the columns.
- Mixed encodings. A file stitched together from several exports can contain chunks in different encodings. The sign is readable text mixed with garbled characters.
- Line breaks inside fields. Addresses and comments often contain newlines. Parsing without regard to quotes turns one record into two.
FAQ
How many rows does Tabulens open?
The limit is 50 million rows per file. There is no limit on file size: only the visible records are read.
Can I export part of a large CSV to Excel?
Yes: apply a filter and export the result to XLSX. Mind the format's own limit of 1,048,575 data rows.
Will a 10 GB file open?
Yes, but building the row index needs one full read of the file — in practice that is disk time, around a minute on a fast SSD.
Free for personal use. Your file is not uploaded to a server. Windows version — 3.5 MB, no installation: details. Organizations: license.