qio reads and writes Apache Parquet files from R.
Read a whole file with read_parquet(), or open a larger one
with open_parquet() to inspect its schema and read only the
columns, row groups, or batches you need. It is built on the bundled C
library carquet and
has no required R package dependencies.
qio is not on CRAN yet. Install the development version from GitHub:
pak::pak("pedrobtz/qio")Building from source needs GNU make and a C compiler. Zstandard and LZ4 are bundled; zlib comes from the system, or from Rtools on Windows.
library(qio)write_parquet() and read_parquet() handle a
whole file at a time.
write_parquet(mtcars, "mtcars.parquet")
cars <- read_parquet("mtcars.parquet")Read part of a file by naming columns or row groups. A column you do not select is never decompressed, which is the cheapest speed-up available on a wide file.
read_parquet("mtcars.parquet", columns = c("mpg", "cyl"))open_parquet() returns a handle. Inspecting one is cheap
because it reads the footer, not the data, so you can look before
deciding what to read.
# several row groups, so there is something to select
write_parquet(mtcars, "mtcars.parquet", row_group_size = 16)
pf <- open_parquet("mtcars.parquet")
pf
#> <qio_parquet_file>
#> mtcars.parquet
#> 32 rows x 11 columns; 2 row groups
schema(pf) # column paths, Parquet types, nullability
row_groups(pf) # rows and bytes per group
read_plan(pf) # the R type each column will become
collect(pf, columns = c("mpg", "cyl"))
collect(pf, row_groups = 1)
close_parquet(pf)Use walk_batches() for a file that does not fit in
memory. It calls your function once per batch and keeps only one batch
alive at a time.
pf <- open_parquet("big.parquet")
walk_batches(pf, batch_size = 100000, FUN = function(batch, index) {
# one data frame at a time
})
close_parquet(pf)Handles hold an open file, so close them when you are done.
read_parquet() opens and closes one for you.
Parquet describes storage with a physical type and meaning with an optional logical type. qio reads them like this by default:
| Parquet | Logical type | R |
|---|---|---|
BOOLEAN |
logical |
|
INT32 |
integer |
|
INT32 |
DATE |
Date |
INT32, INT64 |
TIME |
double (seconds) |
INT64 |
double |
|
INT64 |
TIMESTAMP |
POSIXct (UTC) |
INT96 |
POSIXct (UTC) |
|
FLOAT, DOUBLE |
double |
|
BYTE_ARRAY |
STRING, ENUM, JSON |
character |
BYTE_ARRAY |
list of raw |
|
FIXED_LEN_BYTE_ARRAY |
list of raw |
|
FIXED_LEN_BYTE_ARRAY |
UUID |
character |
FIXED_LEN_BYTE_ARRAY |
FLOAT16 |
double |
| any | DECIMAL |
double |
Bytes are text only when the file says so, which is why an
unannotated BYTE_ARRAY stays raw. INT64 is
exact through 2^53 and NA beyond it; pass
int64 = "integer64" for the full range. Nulls become
NA. Nested and repeated columns are skipped for now.
Writing infers the reverse, and parquet_schema()
overrides it per column:
| R | Parquet |
|---|---|
logical |
BOOLEAN |
integer |
INT32 |
double |
DOUBLE |
character, factor |
BYTE_ARRAY + STRING |
Date |
INT32 + DATE |
POSIXct |
INT64 + TIMESTAMP (UTC, microseconds) |
types <- parquet_schema(mpg = "FLOAT", cyl = "INT64")
write_parquet(mtcars, "mtcars.parquet", schema = types)read_plan() reports what any file will produce before
you read it, and ?qio-types documents every mapping,
including where precision is lost.
write_parquet() takes compression,
row_group_size, metadata,
sorted_by, and append. Beyond
schema() and row_groups(), a handle can report
column_chunks(), column_statistics(),
page_index(), metadata(), and
bloom_filter_may_contain(). validate_parquet()
checks that a file is structurally sound without reading it.
See ?qio-limitations for what qio deliberately does not
do.
qio is MIT licensed. It bundles third-party C sources, each under its own license and shipped with its license file:
| Bundled | License | Location |
|---|---|---|
| carquet | MIT | src/carquet/LICENSE |
| Snappy (in carquet) | BSD-3-Clause | src/carquet/compression/snappy.c |
| Zstandard | BSD-3-Clause | src/zstd/LICENSE |
| LZ4 | BSD-2-Clause | src/lz4/LICENSE |
inst/COPYRIGHTS lists every copyright holder, the files
each covers, and the modifications qio makes; DESCRIPTION
points at it through its Copyright field. Each bundled
library is pinned to an exact upstream commit.
qio carries local patches to carquet; most fix defects that silently
corrupted or rejected valid data. They live as individual commits on the
qio branch of a carquet
fork, which is what src/carquet is vendored from, so
each one can be read on its own and offered upstream. The reason for
each is recorded in the repository’s .agents/VENDORED.md,
which is not shipped in the source package.