Why large XML files hang everything
A 200 MB file takes ten minutes to open in an editor and eats three gigabytes of memory. The editor is not bad: it is the default way XML gets parsed that is expensive.
Tree versus stream
There are two ways to parse XML.
Build a tree. The whole document is turned into objects: an element, its attributes, its children, a link to its parent. It is convenient — you can walk anywhere — and expensive: every element carries overhead several times larger than the text itself. Rule of thumb: the tree weighs 5–10 times more than the file.
Read as a stream. Parsing goes sequentially and reports events: “element opened”, “text”, “element closed”. Memory is spent only on what you decided to keep. It is less convenient — you cannot go back — but the file size stops being a problem.
A browser builds a tree when you open an .xml file. Excel does the same on import. That explains how both behave on big files.
How we do it
Parsing takes two passes over the text, with no tree at all.
The first pass only counts: which elements occur, at what depth, how many times, and how many different child elements they contain. Memory goes into a small dictionary of paths — tens of kilobytes regardless of file size. The result is a list of candidates for “the table row”, with counts.
The second pass extracts only the element you chose. Memory grows with the resulting table, not with the document: from a 200 MB file where you need 50,000 orders with 15 fields each, you get a table of a few tens of megabytes.
An honest limit: all this happens in the browser, so the document text is still read in full. Above 512 MB we refuse to open the file with a clear message, instead of hanging the tab.
What to do with files over a gigabyte
- Split by period or account on the side of the system that produces the file. This is usually possible and useful in itself.
- Convert to CSV with a streaming script and work with that: a CSV is limited by nothing but your disk.
- Load it into a database. If volumes like this arrive regularly, a viewer is not the solution you need.
Small things that trip up parsers
- A
>character inside an attribute value. You cannot find the end of a tag with a plain search — quotes have to be taken into account. - CDATA. There is no markup inside such blocks; everything is text, even if it looks like tags.
- Self-closing elements.
<item/>is an opening and a closing at once. A mistake in handling it shifts all the following nesting, and the table comes out wrong without any sign of it. - Entities.
&and numeric references such asЯmust be expanded, otherwise markup is left in your data.
FAQ
How much memory does a 100 MB file need?
With streaming parsing, roughly the size of the resulting table plus the document text itself. With a tree, from 500 MB to a gigabyte.
Why does Notepad++ open it while the browser hangs?
An editor shows text and does not parse the markup. A browser builds a tree and applies styles to it.
Can I look at just the first thousand records?
Yes, and it is a sensible way to understand the structure: choose the row element and apply a filter.
Free for personal use. Your file is not uploaded to a server. Windows version — 3.5 MB, no installation: details. Organizations: license.