Skip to contents

Read, time-travel and append to Apache Iceberg tables, straight from R.
No Spark, no JVM, no SQL engine in the middle. 🧊

πŸ“¦ Installation

# the CRAN release
install.packages("icebergr")

# or the development version, from r-universe
install.packages("icebergr", repos = c(
  "https://pursuitofdatascience.r-universe.dev",
  "https://cloud.r-project.org"
))

On Windows and macOS that is a prebuilt binary. On Linux, R compiles it, so install Rust 1.88 or newer first:

curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh

πŸ” Read

library(icebergr)
tbl <- icebergr_example_table() # a real Iceberg table, built on your machine

icebergr_collect(icebergr_scan(tbl, filter = amount > 997, select = c("id", "event", "amount")))
#> # A tibble: 3 Γ— 3
#>      id event    amount
#>   <int> <chr>     <dbl>
#> 1  1498 refund      998
#> 2  1499 purchase    999
#> 3  1500 refund     1000

The filter runs inside Iceberg, which skips whole files before reading a byte:

A filter written in R is pushed down into iceberg-rust, which plans the scan, prunes files and row groups, reads two of six data files, and returns the result to R over the Arrow C stream.

⏳ Time travel and ✍️ appends

first <- icebergr_snapshots(tbl)$snapshot_id[[1]]
nrow(icebergr_collect(icebergr_scan(tbl, snapshot_id = first)))
#> [1] 500

tbl <- icebergr_append(tbl, head(icebergr_collect(tbl), 2)) # a new snapshot, nothing rewritten
nrow(icebergr_collect(tbl))
#> [1] 1002

πŸ”Œ Your own catalog

Put your token in ~/.Renviron as ICEBERGR_REST_TOKEN=... and restart R, then:

catalog <- icebergr_catalog("rest", uri = "https://your.catalog")
tbl <- icebergr_table(catalog, "db.events")

AWS Glue, S3 and every other setting: catalog configuration.

πŸ—ΊοΈ What works

βœ… Works 🚫 Not yet
Read filter and column pushdown Β· merge-on-read deletes Β· struct, list, map πŸ¦€ limit pushdown Β· πŸ¦€ filters on nested fields
Time travel by snapshot id or time Β· the schema as of any snapshot
Write append Β· create a table or namespace Β· register a table 🚧 partitioned appends Β· πŸ¦€ deletes, MERGE, overwrite
Catalogs REST Β· local memory Β· βš™οΈ AWS Glue Β· βš™οΈ S3 πŸ¦€ Hadoop (use memory)

βš™οΈ opt-in at build time Β· πŸ¦€ missing in iceberg-rust itself Β· 🚧 not in this release Β· icebergr_spec_support() lists everything for your build.

⚠️ Good to know

πŸ”‘ Credentials come from environment variables, never arguments, and never travel over plain http://.
πŸ†” Snapshot ids are character: they are 64-bit integers, and a double holds 53 bits.
πŸ”’ Install bit64 so long columns stay exact past 2^53.
🧡 Parallel work: parallel::makeCluster(), not mclapply(), and connect inside each worker.
🐒 limit applies after the scan. To read less, filter.

πŸ“œ Licence

GPL (>= 3); bundled Rust crates keep their own licences (inst/NOTICE, LICENSE.note). Apache, Apache Iceberg and Iceberg are trademarks of The Apache Software Foundation; icebergr is a community package, not affiliated with or endorsed by the ASF.