Skip to contents

Connection recipes for the catalogs icebergr supports, how credentials are handled, and what to do when a connection does not work. Code in this vignette is shown but not run.

Credentials are never arguments

icebergr reads credentials from environment variables and provides no argument to pass one. This is deliberate. A token passed as an argument ends up in the script that called it, in .Rhistory, in any knitr cache of the chunk that ran it, and in the traceback() of any error raised nearby.

Variable Iceberg property Used for
ICEBERGR_REST_TOKEN token Bearer token for a REST catalog
ICEBERGR_REST_CREDENTIAL credential OAuth2 client credential
ICEBERGR_REST_OAUTH2_SERVER_URI oauth2-server-uri OAuth2 token endpoint
ICEBERGR_REST_SCOPE scope OAuth2 scope
ICEBERGR_S3_ACCESS_KEY_ID s3.access-key-id Object storage access key
ICEBERGR_S3_SECRET_ACCESS_KEY s3.secret-access-key Object storage secret
ICEBERGR_S3_SESSION_TOKEN s3.session-token Temporary session token
ICEBERGR_S3_REGION s3.region Object storage region
ICEBERGR_S3_ENDPOINT s3.endpoint Custom or S3-compatible endpoint

The standard AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, AWS_SESSION_TOKEN and AWS_REGION variables are consulted as a fallback, so an environment already configured for AWS works without further setup, but only for a connection that addresses object storage: storage = "s3", type = "glue", or an s3:// warehouse. They are not forwarded to a catalog with no object storage in sight, because a REST catalog controls each table’s location and may answer with its own s3.endpoint, at which point ambient keys would sign requests to a host it chose. If a REST catalog identified by name needs them, say storage = "s3".

Nothing goes out unencrypted, either. Whenever any credential property is populated, an http:// uri or OAuth2 endpoint is an error rather than a request carrying Authorization: Bearer … in the clear. A loopback address is exempt, since developing against a catalog on your own machine is ordinary, and ICEBERGR_ALLOW_INSECURE_CREDENTIALS=true overrides the check for the case where you know the network is trusted.

Set them outside your scripts: in ~/.Renviron, in your shell profile, or from your platform’s secret manager.

# ~/.Renviron
ICEBERGR_REST_TOKEN=eyJhbGciOi...
ICEBERGR_S3_REGION=us-east-1

Two guarantees worth relying on:

  • print() on a catalog never shows its properties.
  • Errors raised while connecting list property keys only, never values.

You can still pass a credential through ... if you must, but icebergr warns when you do, because it is nearly always a mistake.

Iceberg REST catalog

The common case: Polaris, Lakekeeper, Nessie, Unity Catalog’s Iceberg endpoint, Tabular-style services, or a self-hosted REST catalog.

Sys.setenv(ICEBERGR_REST_TOKEN = "...") # better: set this in ~/.Renviron

catalog <- icebergr_catalog(
  "rest",
  uri = "https://catalog.example.com/api/catalog",
  warehouse = "analytics"
)

icebergr_list_namespaces(catalog)
icebergr_list_tables(catalog, "analytics")

warehouse is whatever the server expects to identify the warehouse, often a name rather than a path.

With OAuth2 client credentials

Sys.setenv(
  ICEBERGR_REST_CREDENTIAL = "client_id:client_secret",
  ICEBERGR_REST_OAUTH2_SERVER_URI = "https://auth.example.com/oauth/token",
  ICEBERGR_REST_SCOPE = "catalog"
)

catalog <- icebergr_catalog("rest", uri = "https://catalog.example.com")

Extra, non-secret properties

Anything iceberg-rust accepts can be passed through .... Its REST client reads a URL prefix, OAuth2’s audience and resource, and sends each header.<name> property as an HTTP header:

catalog <- icebergr_catalog(
  "rest",
  uri = "https://catalog.example.com",
  prefix = "analytics",
  `header.X-Iceberg-Access-Delegation` = "vended-credentials"
)

A property it does not know is ignored without a word, so check the spelling. There is no SigV4 request signing in iceberg-rust 0.10, so a REST endpoint that requires it, such as AWS Glue’s, is out of reach with type = "rest"; use type = "glue", which goes through the AWS SDK, instead.

Object storage

REST catalogs usually serve tables that live in S3. That needs the optional s3 Cargo feature, which is off by default because opendal is a substantial subtree:

ICEBERGR_CARGO_FEATURES=s3 R CMD INSTALL --preclean .

That build needs network access. The crates vendored inside the CRAN tarball cover the default install only, so cargo fetches the s3 subtree from crates.io.

Check before relying on it:

"s3" %in% icebergr_spec_support()$cargo_features
Sys.setenv(
  ICEBERGR_S3_ACCESS_KEY_ID = "...",
  ICEBERGR_S3_SECRET_ACCESS_KEY = "...",
  ICEBERGR_S3_REGION = "us-east-1"
)

catalog <- icebergr_catalog(
  "rest",
  uri = "https://catalog.example.com",
  storage = "s3"
)

S3-compatible storage (MinIO, R2, Ceph)

Sys.setenv(
  ICEBERGR_S3_ENDPOINT = "https://minio.internal:9000",
  ICEBERGR_S3_ACCESS_KEY_ID = "...",
  ICEBERGR_S3_SECRET_ACCESS_KEY = "..."
)

catalog <- icebergr_catalog(
  "rest",
  uri = "https://catalog.internal/api/catalog",
  storage = "s3",
  `s3.path-style-access` = "true"
)

Most S3-compatible services need path-style access.

AWS Glue

Needs the glue Cargo feature, which implies s3:

ICEBERGR_CARGO_FEATURES=glue R CMD INSTALL --preclean .
catalog <- icebergr_catalog(
  "glue",
  warehouse = "s3://my-bucket/warehouse",
  region_name = "us-east-1"
)

icebergr_list_namespaces(catalog)

Glue uses the AWS SDK’s own credential chain, so instance roles, IRSA, SSO profiles and AWS_PROFILE all work as they do elsewhere. region_name is the property the Glue client reads for its region; without it, the SDK’s usual sources decide, such as AWS_REGION or the profile.

If the feature is not compiled in, the error says so and how to get it, rather than failing at link time.

Local warehouses

For local files, testing, or a warehouse on a mounted filesystem, use the in-process memory catalog. This needs no server and no credentials, and is what this package’s own tests run against.

warehouse <- "/data/warehouse"
catalog <- icebergr_catalog("memory", warehouse = warehouse)

icebergr_create_namespace(catalog, "db")
tbl <- icebergr_create_table(catalog, "db.events", data.frame(id = integer()))

There is deliberately no "hadoop" catalog type: iceberg-rust does not implement a Hadoop or filesystem catalog. memory is the local equivalent.

Its one real limitation is that the table registry lives in memory, so it does not persist between sessions. Tables already on disk are re-attached with icebergr_register_table():

catalog <- icebergr_catalog("memory", warehouse = "/data/warehouse")
icebergr_create_namespace(catalog, "db")

tbl <- icebergr_register_table(
  catalog, "db.events",
  "/data/warehouse/db/events/metadata/00003-....metadata.json"
)

Point it at the newest metadata file: Iceberg writes a new one per commit, and the newest is the current state of the table.

The file has to be inside the catalog’s own warehouse. That is confine = TRUE, the default, and it matters because a metadata file names absolute paths for its location, its manifests and every data file, so registering one from a shared drive or an issue attachment reads whatever its author nominated, and with the s3 feature compiled in can make an outbound request from a catalog you opened offline. confine = FALSE lifts the restriction for a file you trust; it does not vet the paths inside the file, which nothing can.

Performance notes

  • Push filters down. filter and select are the difference between reading a table and reading the part of it you want. Confirm with icebergr_scan_plan().

  • limit is not pushdown. iceberg-rust has no row limit in its scan API, so limit bounds decoding, not planning. To read less, filter.

  • Batch size trades memory for call overhead. batch_size controls rows per Arrow batch; the default is usually right.

  • Worker threads. The tokio runtime uses 2 worker threads, which keeps parallelism within CRAN’s limits for checks. For large scans over object storage, raise it before loading the package:

    Sys.setenv(ICEBERGR_WORKER_THREADS = "8")
    library(icebergr)

    The runtime starts on first use, so this has no effect once you have connected.

  • Parallelise with a PSOCK cluster, not with fork. icebergr drives an async runtime whose worker threads do not survive fork(), and R’s parallel::mclapply() forks. Once you have made any icebergr call in the parent, a forked child’s calls into icebergr fail with an error that says so. Use parallel::makeCluster(), which starts fresh R processes, and open the catalog inside each worker:

    cl <- parallel::makeCluster(4)
    parallel::clusterEvalQ(cl, library(icebergr))
    parallel::parLapply(cl, shards, function(shard) {
      catalog <- icebergr_catalog("memory", warehouse = "/data/warehouse")
      tbl <- icebergr_table(catalog, "db.events")
      icebergr_collect(icebergr_scan(tbl, filter = day == shard))
    })

    Note that the handle is opened in the worker. A handle cannot be sent to one: it holds a pointer into the Rust side, and serialising it produces the “no longer usable” error rather than a wrong answer.

  • Reads stream. Batches are pulled one at a time rather than collected up front, so memory stays bounded by batch size, not table size, until icebergr_collect() materialises the result into a tibble, which is necessarily all of it. To process a table larger than memory, scan it in filtered pieces.

Troubleshooting

icebergr_catalog() succeeded but the next call failed. Building a REST or Glue handle does not contact the server; iceberg-rust connects on the first operation that needs it. A mistyped uri, an unreachable host or an unset token therefore surface at icebergr_list_namespaces() or icebergr_table() rather than at the call that looks like it should have caught them. Ask the catalog something immediately if you want to find out then and there.

could not connect to the REST catalog, with a list of property keys. The keys tell you what the catalog was given. A missing token usually means the environment variable is not visible to R: check Sys.getenv("ICEBERGR_REST_TOKEN"), and remember that ~/.Renviron is read only at startup.

TLS failures behind a corporate proxy. icebergr uses rustls with the platform trust store, so a CA installed in the system store is honoured. A CA present only in R’s own bundle is not.

is not available in this build of icebergr. An optional Cargo feature is missing. The message names the feature and gives the install command.

The compiled Rust component of icebergr is not available. The package loaded but its native routines did not. This is a half-finished source install; reinstall and read the build log for cargo errors.

Reads succeed, writes fail with a permissions error. Iceberg commits write both data files and new metadata. Write access to the table’s metadata/ prefix is required, not just to data/.

A table shows fewer rows than expected. Reading a merge-on-read table is supported: iceberg-rust applies positional and equality deletes during the scan, so rows another engine deleted will correctly be absent. If the count still looks wrong, check whether you are reading an older snapshot (print() on a table shows which one) and whether a filter is pruning more than you meant, with icebergr_scan_plan().

Writes to a table another engine performs deletes on. icebergr only appends, which is always safe. It cannot delete or update rows, so a workflow needing that must do it elsewhere; see icebergr_spec_support().

this process was forked from one that had already used icebergr. The async runtime’s worker threads do not survive a fork(), which is what parallel::mclapply() does, so a forked child cannot complete a call once the parent has started the runtime. It refuses at once rather than waiting forever on threads that did not come across. Forking before any icebergr call is fine, since each child then starts its own runtime, but once the parent has connected, use parallel::makeCluster() instead and open the catalog inside each worker. See Performance notes above.

cannot append to "db.events": it is partitioned by …. Reading a partitioned table works, and icebergr_partitions() reports its spec, but appending to one would have to compute a partition value for every row, which this version does not do. The append is refused before any data file is written, so nothing is left behind in the warehouse. Write to such a table with an engine that supports partitioned writes.

Writing to a partitioned table is one of several operations Iceberg’s spec defines (The Apache Software Foundation 2026) that iceberg-rust 0.10 does not yet implement; vignette("writing") lists the rest and says which side each gap is on.

References

The Apache Software Foundation. 2026. Apache Iceberg Table Spec. Apache Iceberg documentation. https://iceberg.apache.org/spec/.