Catalog configuration
Youzhi
Yu
University of Chicago
Source: vignettes/catalog-configuration.Rmd
catalog-configuration.RmdConnection recipes for the catalogs icebergr supports,
how credentials are handled, and what to do when a connection does not
work. Code in this vignette is shown but not run.
Credentials are never arguments
icebergr reads credentials from environment variables
and provides no argument to pass one. This is deliberate. A token passed
as an argument ends up in the script that called it, in
.Rhistory, in any knitr cache of the chunk that ran it, and
in the traceback() of any error raised nearby.
| Variable | Iceberg property | Used for |
|---|---|---|
ICEBERGR_REST_TOKEN |
token |
Bearer token for a REST catalog |
ICEBERGR_REST_CREDENTIAL |
credential |
OAuth2 client credential |
ICEBERGR_REST_OAUTH2_SERVER_URI |
oauth2-server-uri |
OAuth2 token endpoint |
ICEBERGR_REST_SCOPE |
scope |
OAuth2 scope |
ICEBERGR_S3_ACCESS_KEY_ID |
s3.access-key-id |
Object storage access key |
ICEBERGR_S3_SECRET_ACCESS_KEY |
s3.secret-access-key |
Object storage secret |
ICEBERGR_S3_SESSION_TOKEN |
s3.session-token |
Temporary session token |
ICEBERGR_S3_REGION |
s3.region |
Object storage region |
ICEBERGR_S3_ENDPOINT |
s3.endpoint |
Custom or S3-compatible endpoint |
The standard AWS_ACCESS_KEY_ID,
AWS_SECRET_ACCESS_KEY, AWS_SESSION_TOKEN and
AWS_REGION variables are consulted as a fallback, so an
environment already configured for AWS works without further setup, but
only for a connection that addresses object storage:
storage = "s3", type = "glue", or an
s3:// warehouse. They are not forwarded to a catalog with
no object storage in sight, because a REST catalog controls each table’s
location and may answer with its own
s3.endpoint, at which point ambient keys would sign
requests to a host it chose. If a REST catalog identified by name needs
them, say storage = "s3".
Nothing goes out unencrypted, either. Whenever any credential
property is populated, an http:// uri or
OAuth2 endpoint is an error rather than a request carrying
Authorization: Bearer … in the clear. A loopback address is
exempt, since developing against a catalog on your own machine is
ordinary, and ICEBERGR_ALLOW_INSECURE_CREDENTIALS=true
overrides the check for the case where you know the network is
trusted.
Set them outside your scripts: in ~/.Renviron, in your
shell profile, or from your platform’s secret manager.
# ~/.Renviron
ICEBERGR_REST_TOKEN=eyJhbGciOi...
ICEBERGR_S3_REGION=us-east-1
Two guarantees worth relying on:
-
print()on a catalog never shows its properties. - Errors raised while connecting list property keys only, never values.
You can still pass a credential through ... if you must,
but icebergr warns when you do, because it is nearly always
a mistake.
Iceberg REST catalog
The common case: Polaris, Lakekeeper, Nessie, Unity Catalog’s Iceberg endpoint, Tabular-style services, or a self-hosted REST catalog.
Sys.setenv(ICEBERGR_REST_TOKEN = "...") # better: set this in ~/.Renviron
catalog <- icebergr_catalog(
"rest",
uri = "https://catalog.example.com/api/catalog",
warehouse = "analytics"
)
icebergr_list_namespaces(catalog)
icebergr_list_tables(catalog, "analytics")warehouse is whatever the server expects to identify the
warehouse, often a name rather than a path.
With OAuth2 client credentials
Sys.setenv(
ICEBERGR_REST_CREDENTIAL = "client_id:client_secret",
ICEBERGR_REST_OAUTH2_SERVER_URI = "https://auth.example.com/oauth/token",
ICEBERGR_REST_SCOPE = "catalog"
)
catalog <- icebergr_catalog("rest", uri = "https://catalog.example.com")Extra, non-secret properties
Anything iceberg-rust accepts can be passed through
.... Its REST client reads a URL prefix,
OAuth2’s audience and resource, and sends each
header.<name> property as an HTTP header:
catalog <- icebergr_catalog(
"rest",
uri = "https://catalog.example.com",
prefix = "analytics",
`header.X-Iceberg-Access-Delegation` = "vended-credentials"
)A property it does not know is ignored without a word, so check the
spelling. There is no SigV4 request signing in iceberg-rust
0.10, so a REST endpoint that requires it, such as AWS Glue’s, is out of
reach with type = "rest"; use type = "glue",
which goes through the AWS SDK, instead.
Object storage
REST catalogs usually serve tables that live in S3. That needs the
optional s3 Cargo feature, which is off by default because
opendal is a substantial subtree:
That build needs network access. The crates vendored inside the CRAN
tarball cover the default install only, so cargo fetches the
s3 subtree from crates.io.
Check before relying on it:
"s3" %in% icebergr_spec_support()$cargo_features
Sys.setenv(
ICEBERGR_S3_ACCESS_KEY_ID = "...",
ICEBERGR_S3_SECRET_ACCESS_KEY = "...",
ICEBERGR_S3_REGION = "us-east-1"
)
catalog <- icebergr_catalog(
"rest",
uri = "https://catalog.example.com",
storage = "s3"
)S3-compatible storage (MinIO, R2, Ceph)
Sys.setenv(
ICEBERGR_S3_ENDPOINT = "https://minio.internal:9000",
ICEBERGR_S3_ACCESS_KEY_ID = "...",
ICEBERGR_S3_SECRET_ACCESS_KEY = "..."
)
catalog <- icebergr_catalog(
"rest",
uri = "https://catalog.internal/api/catalog",
storage = "s3",
`s3.path-style-access` = "true"
)Most S3-compatible services need path-style access.
AWS Glue
Needs the glue Cargo feature, which implies
s3:
catalog <- icebergr_catalog(
"glue",
warehouse = "s3://my-bucket/warehouse",
region_name = "us-east-1"
)
icebergr_list_namespaces(catalog)Glue uses the AWS SDK’s own credential chain, so instance roles,
IRSA, SSO profiles and AWS_PROFILE all work as they do
elsewhere. region_name is the property the Glue client
reads for its region; without it, the SDK’s usual sources decide, such
as AWS_REGION or the profile.
If the feature is not compiled in, the error says so and how to get it, rather than failing at link time.
Local warehouses
For local files, testing, or a warehouse on a mounted filesystem, use
the in-process memory catalog. This needs no server and no
credentials, and is what this package’s own tests run against.
warehouse <- "/data/warehouse"
catalog <- icebergr_catalog("memory", warehouse = warehouse)
icebergr_create_namespace(catalog, "db")
tbl <- icebergr_create_table(catalog, "db.events", data.frame(id = integer()))There is deliberately no "hadoop" catalog type:
iceberg-rust does not implement a Hadoop or
filesystem catalog. memory is the local
equivalent.
Its one real limitation is that the table registry lives in memory,
so it does not persist between sessions. Tables already on disk are
re-attached with icebergr_register_table():
catalog <- icebergr_catalog("memory", warehouse = "/data/warehouse")
icebergr_create_namespace(catalog, "db")
tbl <- icebergr_register_table(
catalog, "db.events",
"/data/warehouse/db/events/metadata/00003-....metadata.json"
)Point it at the newest metadata file: Iceberg writes a new one per commit, and the newest is the current state of the table.
The file has to be inside the catalog’s own warehouse. That is
confine = TRUE, the default, and it matters because a
metadata file names absolute paths for its location, its
manifests and every data file, so registering one from a shared drive or
an issue attachment reads whatever its author nominated, and with the
s3 feature compiled in can make an outbound request from a
catalog you opened offline. confine = FALSE lifts the
restriction for a file you trust; it does not vet the paths
inside the file, which nothing can.
Performance notes
Push filters down.
filterandselectare the difference between reading a table and reading the part of it you want. Confirm withicebergr_scan_plan().limitis not pushdown.iceberg-rusthas no row limit in its scan API, solimitbounds decoding, not planning. To read less, filter.Batch size trades memory for call overhead.
batch_sizecontrols rows per Arrow batch; the default is usually right.-
Worker threads. The tokio runtime uses 2 worker threads, which keeps parallelism within CRAN’s limits for checks. For large scans over object storage, raise it before loading the package:
Sys.setenv(ICEBERGR_WORKER_THREADS = "8") library(icebergr)The runtime starts on first use, so this has no effect once you have connected.
-
Parallelise with a PSOCK cluster, not with
fork.icebergrdrives an async runtime whose worker threads do not survivefork(), and R’sparallel::mclapply()forks. Once you have made anyicebergrcall in the parent, a forked child’s calls intoicebergrfail with an error that says so. Useparallel::makeCluster(), which starts fresh R processes, and open the catalog inside each worker:cl <- parallel::makeCluster(4) parallel::clusterEvalQ(cl, library(icebergr)) parallel::parLapply(cl, shards, function(shard) { catalog <- icebergr_catalog("memory", warehouse = "/data/warehouse") tbl <- icebergr_table(catalog, "db.events") icebergr_collect(icebergr_scan(tbl, filter = day == shard)) })Note that the handle is opened in the worker. A handle cannot be sent to one: it holds a pointer into the Rust side, and serialising it produces the “no longer usable” error rather than a wrong answer.
Reads stream. Batches are pulled one at a time rather than collected up front, so memory stays bounded by batch size, not table size, until
icebergr_collect()materialises the result into a tibble, which is necessarily all of it. To process a table larger than memory, scan it in filtered pieces.
Troubleshooting
icebergr_catalog() succeeded but the next call
failed. Building a REST or Glue handle does not contact the
server; iceberg-rust connects on the first operation that
needs it. A mistyped uri, an unreachable host or an unset
token therefore surface at icebergr_list_namespaces() or
icebergr_table() rather than at the call that looks like it
should have caught them. Ask the catalog something immediately if you
want to find out then and there.
could not connect to the REST catalog, with a
list of property keys. The keys tell you what the catalog was
given. A missing token usually means the environment
variable is not visible to R: check
Sys.getenv("ICEBERGR_REST_TOKEN"), and remember that
~/.Renviron is read only at startup.
TLS failures behind a corporate proxy.
icebergr uses rustls with the platform trust store, so a CA
installed in the system store is honoured. A CA present only in R’s own
bundle is not.
is not available in this build of icebergr.
An optional Cargo feature is missing. The message names the feature and
gives the install command.
The compiled Rust component of icebergr is not available.
The package loaded but its native routines did not. This is a
half-finished source install; reinstall and read the build log for cargo
errors.
Reads succeed, writes fail with a permissions error.
Iceberg commits write both data files and new metadata. Write access to
the table’s metadata/ prefix is required, not just to
data/.
A table shows fewer rows than expected. Reading a
merge-on-read table is supported: iceberg-rust applies
positional and equality deletes during the scan, so rows another engine
deleted will correctly be absent. If the count still looks wrong, check
whether you are reading an older snapshot (print() on a
table shows which one) and whether a filter is pruning more than you
meant, with icebergr_scan_plan().
Writes to a table another engine performs deletes
on. icebergr only appends, which is always safe.
It cannot delete or update rows, so a workflow needing that must do it
elsewhere; see icebergr_spec_support().
this process was forked from one that had already used icebergr.
The async runtime’s worker threads do not survive a fork(),
which is what parallel::mclapply() does, so a forked child
cannot complete a call once the parent has started the runtime. It
refuses at once rather than waiting forever on threads that did not come
across. Forking before any icebergr call is fine,
since each child then starts its own runtime, but once the parent has
connected, use parallel::makeCluster() instead and open the
catalog inside each worker. See Performance notes above.
cannot append to "db.events": it is partitioned by ….
Reading a partitioned table works, and
icebergr_partitions() reports its spec, but appending to
one would have to compute a partition value for every row, which this
version does not do. The append is refused before any data file is
written, so nothing is left behind in the warehouse. Write to such a
table with an engine that supports partitioned writes.
Writing to a partitioned table is one of several operations Iceberg’s
spec defines (The Apache Software Foundation
2026) that iceberg-rust 0.10 does not yet implement;
vignette("writing") lists the rest and says which side each
gap is on.