Read, time-travel and append to Apache Iceberg tables, straight from R.
No Spark, no JVM, no SQL engine in the middle. π§
π¦ Installation
# the CRAN release
install.packages("icebergr")
# or the development version, from r-universe
install.packages("icebergr", repos = c(
"https://pursuitofdatascience.r-universe.dev",
"https://cloud.r-project.org"
))On Windows and macOS that is a prebuilt binary. On Linux, R compiles it, so install Rust 1.88 or newer first:
π Read
library(icebergr)
tbl <- icebergr_example_table() # a real Iceberg table, built on your machine
icebergr_collect(icebergr_scan(tbl, filter = amount > 997, select = c("id", "event", "amount")))
#> # A tibble: 3 Γ 3
#> id event amount
#> <int> <chr> <dbl>
#> 1 1498 refund 998
#> 2 1499 purchase 999
#> 3 1500 refund 1000The filter runs inside Iceberg, which skips whole files before reading a byte:

β³ Time travel and βοΈ appends
first <- icebergr_snapshots(tbl)$snapshot_id[[1]]
nrow(icebergr_collect(icebergr_scan(tbl, snapshot_id = first)))
#> [1] 500
tbl <- icebergr_append(tbl, head(icebergr_collect(tbl), 2)) # a new snapshot, nothing rewritten
nrow(icebergr_collect(tbl))
#> [1] 1002π Your own catalog
Put your token in ~/.Renviron as ICEBERGR_REST_TOKEN=... and restart R, then:
catalog <- icebergr_catalog("rest", uri = "https://your.catalog")
tbl <- icebergr_table(catalog, "db.events")AWS Glue, S3 and every other setting: catalog configuration.
πΊοΈ What works
| β Works | π« Not yet | |
|---|---|---|
| Read | filter and column pushdown Β· merge-on-read deletes Β· struct, list, map
|
π¦ limit pushdown Β· π¦ filters on nested fields |
| Time travel | by snapshot id or time Β· the schema as of any snapshot | |
| Write | append Β· create a table or namespace Β· register a table | π§ partitioned appends Β· π¦ deletes, MERGE, overwrite |
| Catalogs | REST Β· local memory Β· βοΈ AWS Glue Β· βοΈ S3 |
π¦ Hadoop (use memory) |
βοΈ opt-in at build time Β· π¦ missing in iceberg-rust itself Β· π§ not in this release Β· icebergr_spec_support() lists everything for your build.
β οΈ Good to know
| π | Credentials come from environment variables, never arguments, and never travel over plain http://. |
| π | Snapshot ids are character: they are 64-bit integers, and a double holds 53 bits. |
| π’ | Install bit64 so long columns stay exact past 2^53. |
| π§΅ | Parallel work: parallel::makeCluster(), not mclapply(), and connect inside each worker. |
| π’ |
limit applies after the scan. To read less, filter. |