Results at a glance
Test machine and files
The test machine runs Linux on an 11th Gen Intel Core i7-11850H with 8 cores, 16 threads, 31.1 GiB of RAM, and an NVMe SSD. We tested a release build of csvlite 0.1.0.
All five files use the same generated 15-column schema, with integer, decimal, date, text, and boolean fields. One percent of the values are invalid and two percent are empty. A fixed seed makes the files reproducible byte for byte.
| File | Bytes | Rows | SHA-256 |
|---|---|---|---|
| 1 MB | 999,837 | 5,452 | 22a2f228a339 |
| 10 MB | 9,999,912 | 53,596 | 6a7e28ed112b |
| 100 MB | 99,999,911 | 532,923 | fb052e3d47fe |
| 1 GB | 999,999,957 | 5,300,205 | 0f8da3f8cec6 |
| 10 GB | 9,999,999,809 | 52,720,416 | 4719e88857cc |
Before every warm-cache trial, the runner reads the entire source file into the page cache. Before every cold-cache trial, it asks Linux to evict that file with POSIX_FADV_DONTNEED. Cache state is controlled per trial rather than inferred from the order of the runs.
Operations
Each cell shows the median warm-cache time followed by the median cold-cache time. Both medians come from ten trials. The small-file figures include process startup, which is why several round to the same millisecond.
| Operation | 1 MB | 10 MB | 100 MB | 1 GB | 10 GB |
|---|---|---|---|---|---|
| Rows | 5,452 | 53,596 | 532,923 | 5,300,205 | 52,720,416 |
| Open and index | 5 / 7 ms | 9 / 13 ms | 17 / 49 ms | 189 / 464 ms | 4.05 / 5.82 s |
| Filter, text equals | 6 / 7 ms | 7 / 14 ms | 23 / 52 ms | 248 / 492 ms | 5.55 / 6.79 s |
| Filter, regular expression | 7 / 9 ms | 13 / 19 ms | 62 / 73 ms | 492 / 760 ms | 7.92 / 8.73 s |
| Filter, numeric range | 5 / 8 ms | 9 / 15 ms | 32 / 62 ms | 358 / 595 ms | 6.70 / 7.73 s |
| Filter, invalid numbers | 6 / 7 ms | 8 / 16 ms | 27 / 56 ms | 293 / 569 ms | 5.77 / 7.59 s |
| Filter, date range | 6 / 7 ms | 9 / 16 ms | 29 / 57 ms | 333 / 564 ms | 5.83 / 7.93 s |
| Filter, invalid dates | 6 / 7 ms | 8 / 17 ms | 25 / 58 ms | 322 / 550 ms | 5.86 / 6.98 s |
| Sort, text column | 7 / 7 ms | 9 / 15 ms | 43 / 68 ms | 455 / 678 ms | 10.57 / 10.25 s |
| Sort, numeric column | 6 / 7 ms | 8 / 15 ms | 52 / 68 ms | 451 / 663 ms | 7.90 / 8.27 s |
| Sort, date column | 6 / 8 ms | 10 / 16 ms | 42 / 68 ms | 416 / 650 ms | 7.93 / 8.21 s |
Range of the 10 GB results
Each cell shows the median followed by the fastest and slowest of the ten trials.
| Operation | Warm: median (min–max) | Cold: median (min–max) |
|---|---|---|
| Open and index | 4.05 s (3.69–5.33) | 5.82 s (5.59–6.57) |
| Filter, text equals | 5.55 s (4.42–7.24) | 6.79 s (5.94–7.70) |
| Filter, regular expression | 7.92 s (7.14–8.44) | 8.73 s (8.24–9.32) |
| Filter, numeric range | 6.70 s (6.39–7.98) | 7.73 s (7.39–8.40) |
| Filter, invalid numbers | 5.77 s (5.45–7.24) | 7.59 s (7.30–7.91) |
| Filter, date range | 5.83 s (5.48–7.08) | 7.93 s (7.18–8.44) |
| Filter, invalid dates | 5.86 s (5.39–6.78) | 6.98 s (6.25–7.72) |
| Sort, text column | 10.57 s (9.86–11.13) | 10.25 s (10.00–10.90) |
| Sort, numeric column | 7.90 s (7.18–8.46) | 8.27 s (7.98–8.54) |
| Sort, date column | 7.93 s (7.25–8.68) | 8.21 s (7.60–8.64) |
| Group by 8 keys | 6.04 s (5.45–7.64) | 7.41 s (6.61–7.80) |
| Group by 56 keys | 6.98 s (6.35–8.40) | 8.37 s (7.49–8.55) |
| Group by 3,654 keys | 7.39 s (6.58–8.11) | 8.76 s (8.06–8.93) |
| Distinct count, 22 values | 7.59 s (7.04–8.68) | 8.88 s (8.54–9.51) |
Warm and cold results overlap in a few CPU-heavy cases, such as the text sort. That is ordinary run-to-run variation: cache control changes how the file is read, but it does not remove scheduling and CPU-frequency noise after the file is open.
Grouping
csvlite calculates grouped results exactly. If a grouping would use too much memory, it stops and reports the limit it reached instead of returning an estimate. Those limits depend on the number of distinct keys and values, not directly on file size.
Number of group keys
Each test calculates a count and a sum. Only the number of distinct group keys changes. As above, times are warm / cold medians and include opening the file.
| Group by | Distinct keys | 1 MB | 10 MB | 100 MB | 1 GB | 10 GB |
|---|---|---|---|---|---|---|
| Region | 8 | 6 / 8 ms | 12 / 16 ms | 34 / 62 ms | 376 / 615 ms | 6.04 / 7.41 s |
| Region + status | 56 | 6 / 8 ms | 11 / 16 ms | 44 / 67 ms | 466 / 684 ms | 6.98 / 8.37 s |
| Date | 3,654 | 10 / 13 ms | 19 / 24 ms | 90 / 106 ms | 684 / 901 ms | 7.39 / 8.76 s |
| Customer e-mail | max 100,005 | 11 / 12 ms | 86 / 88 ms | refused | refused | refused |
| Quantity | max 1,000,005 | 10 / 10 ms | 81 / 86 ms | refused | refused | refused |
csvlite allows up to 100,000 distinct key combinations. In each of the final two rows, the operation and column stay the same as the file grows. The 1 MB and 10 MB tests succeed with 5,452 and 53,596 distinct values; the larger files cross the limit and are refused. The 10 GB tests below the limit finish in roughly six to nine seconds, including the time required to open the file.
Number of values counted
All of these tests group the rows into 56 keys. The only change is the number of distinct values that csvlite must retain for each group. Times are warm / cold medians.
| Distinct count over | Value space | 1 MB | 10 MB | 100 MB | 1 GB | 10 GB |
|---|---|---|---|---|---|---|
| Description | 22 | 7 / 8 ms | 14 / 20 ms | 57 / 79 ms | 579 / 745 ms | 7.59 / 8.88 s |
| Date | 3,654 | 8 / 10 ms | 25 / 31 ms | 183 / 182 ms | refused | refused |
| Customer e-mail | up to 100,005 | 7 / 8 ms | 26 / 30 ms | 369 / 394 ms | refused | refused |
The limit is 2,000,000 retained (group, value) pairs across all worker threads before their results are merged. That total reflects peak memory use. On this 16-thread machine, a result with 205,000 pairs can briefly occupy about 3.3 million pairs across the workers and exceed the limit. The same aggregation can complete with fewer worker threads, so this boundary depends on the number of cores available.
Method
- The published matrix contains 1,800 runs: 18 operations, five file sizes, ten trials, and two cache states.
- Each operation runs in a separate process and opens the file from scratch.
- Sorts build the complete row order in memory. Filters continue until they have a final count of matching rows.
- We checked result counts against an independent implementation before keeping the timings.
- The runner explicitly preloads or evicts the source file before each trial.
- The grouping tests deliberately cross csvlite’s limits. Of the 1,800 runs, 200 were refused and recorded as such.
Tests run 29 August 2026. These numbers measure csvlite only, using the 15-column files and machine described above.