---
description: Learn how to enable the R2 Data Catalog on your bucket, load sample data, and run your first query.
title: Getting started
image: https://developers.cloudflare.com/og-docs.png
---

[Skip to content](#main-content)

> Documentation Index  
> Fetch the complete documentation index at: https://developers.cloudflare.com/r2-data-catalog/llms.txt  
> Use this file to discover all available pages before exploring further.

# Getting started

Last updated Sep 3, 2026|Copy as Markdown| [View as Markdown](https://641a99ee.previews.developers.cloudflare.com/r2-data-catalog/get-started/index.md)| [Agent setup](https://641a99ee.previews.developers.cloudflare.com/agent-setup/)

This guide will instruct you through:

- Creating your first [R2 bucket](https://641a99ee.previews.developers.cloudflare.com/r2/buckets/) and enabling its [data catalog](https://641a99ee.previews.developers.cloudflare.com/r2-data-catalog/).
- Creating an [API token](https://641a99ee.previews.developers.cloudflare.com/r2/api/tokens/) needed for query engines to authenticate with your data catalog.
- Using [PyIceberg ↗](https://py.iceberg.apache.org/) to create your first Iceberg table in a [marimo ↗](https://marimo.io/) Python notebook.
- Using [PyIceberg ↗](https://py.iceberg.apache.org/) to load sample data into your table and query it.

## Prerequisites

1. Sign up for a [Cloudflare account ↗](https://dash.cloudflare.com/sign-up/workers-and-pages).
2. Install [`Node.js` ↗](https://docs.npmjs.com/downloading-and-installing-node-js-and-npm).

<details>

<summary>

Node.js version manager

</summary>

Use a Node version manager like <a href="https://volta.sh/">Volta ↗</a> or <a href="https://github.com/nvm-sh/nvm">nvm ↗</a> to avoid permission issues and change Node.js versions. <a href="https://641a99ee.previews.developers.cloudflare.com/workers/wrangler/install-and-update/">Wrangler</a>, discussed later in this guide, requires a Node version of <code>16.17.0</code> or later.

</details>

## 1. Create an R2 bucket and enable the data catalog

1. If not already logged in, run:

   ```bash
   npx wrangler login
   ```


2. Create an R2 bucket:

   ```bash
   npx wrangler r2 bucket create r2-data-catalog-tutorial
   ```


3. Enable the catalog on your bucket:

   ```bash
   npx wrangler r2 bucket catalog enable r2-data-catalog-tutorial
   ```

   When you run this command, take note of the **Warehouse** and **Catalog URI**. You will need these later.

1. In the Cloudflare dashboard, go to the **R2 Data Catalog** page. [Go to **R2 Data Catalog** ↗](https://dash.cloudflare.com/?to=/:account/data-catalog/overview)
2. Select **Create catalog**.
3. Enter the bucket name `r2-data-catalog-tutorial`. The wizard creates the bucket automatically if it does not already exist. Optionally choose a location hint.
4. Review the configuration and select **Create catalog**.
5. Once created, the catalog detail page displays your **Catalog URI** and **Warehouse name**. Note these values for later.

## 2. Create an API token

Iceberg clients (including [PyIceberg ↗](https://py.iceberg.apache.org/)) must authenticate to the catalog with an [R2 API token](https://641a99ee.previews.developers.cloudflare.com/r2/api/tokens/) that has both R2 and catalog permissions.

1. In the Cloudflare dashboard, go to the **R2 object storage** page. [Go to **Overview** ↗](https://dash.cloudflare.com/?to=/:account/r2/overview)
2. Select **Manage API tokens**.
3. Select **Create API token**.
4. Select the **R2 Token** text to edit your API token name.
5. Under **Permissions**, choose the **Admin Read & Write** permission. This guide creates and writes to tables, so it requires read and write access. For query-only clients, you can instead use an **Admin Read only** token. For details on choosing the right permission level, refer to [Authenticate your Iceberg engine](https://641a99ee.previews.developers.cloudflare.com/r2-data-catalog/manage-catalogs/#authenticate-your-iceberg-engine).
6. Select **Create API Token**.
7. Note the **Token value**.

## 3. Install uv

You need to install a Python package manager. In this guide, use [uv ↗](https://docs.astral.sh/uv/). If you do not already have uv installed, follow the [installing uv guide ↗](https://docs.astral.sh/uv/getting-started/installation/).

## 4. Install marimo and set up your project with uv

We will use [marimo ↗](https://github.com/marimo-team/marimo) as a Python notebook.

1. Create a directory where our notebook will be stored:

   ```bash
   mkdir r2-data-catalog-notebook
   ```


2. Change into our new directory:

   ```bash
   cd r2-data-catalog-notebook
   ```


3. Initialize a new uv project (this creates a `.venv` and a `pyproject.toml`):

   ```plaintext
   uv init
   ```


4. Add marimo and required dependencies:

   ```py
   uv add marimo pyiceberg pyarrow pandas
   ```



## 5. Create a Python notebook to interact with the data warehouse

1. Create a file called `r2-data-catalog-tutorial.py`.
2. Paste the following code snippet into your `r2-data-catalog-tutorial.py` file:

   ```py
   import marimo

   __generated_with = "0.11.31"
   app = marimo.App(width="medium")


   @app.cell
   def _():
   		import marimo as mo
   		return (mo,)


   @app.cell
   def _():
   		import pandas
   		import pyarrow as pa
   		import pyarrow.compute as pc
   		import pyarrow.parquet as pq

   		from pyiceberg.catalog.rest import RestCatalog

   		# Define catalog connection details (replace variables)
   		WAREHOUSE = "<WAREHOUSE>"
   		TOKEN = "<TOKEN>"
   		CATALOG_URI = "<CATALOG_URI>"

   		# Connect to R2 Data Catalog
   		catalog = RestCatalog(
   				name="my_catalog",
   				warehouse=WAREHOUSE,
   				uri=CATALOG_URI,
   				token=TOKEN,
   		)
   		return (
   				CATALOG_URI,
   				RestCatalog,
   				TOKEN,
   				WAREHOUSE,
   				catalog,
   				pa,
   				pandas,
   				pc,
   				pq,
   		)


   @app.cell
   def _(catalog):
   		# Create default namespace if needed
   		catalog.create_namespace_if_not_exists("default")
   		return


   @app.cell
   def _(pa):
   		# Create simple PyArrow table
   		df = pa.table({
   				"id": [1, 2, 3],
   				"name": ["Alice", "Bob", "Charlie"],
   				"score": [80.0, 92.5, 88.0],
   		})
   		return (df,)


   @app.cell
   def _(catalog, df):
   		# Create or load Iceberg table
   		test_table = ("default", "people")
   		if not catalog.table_exists(test_table):
   				print(f"Creating table: {test_table}")
   				table = catalog.create_table(
   						test_table,
   						schema=df.schema,
   				)
   		else:
   				table = catalog.load_table(test_table)
   		return table, test_table


   @app.cell
   def _(df, table):
   		# Append data
   		table.append(df)
   		return


   @app.cell
   def _(table):
   		print("Table contents:")
   		scanned = table.scan().to_arrow()
   		print(scanned.to_pandas())
   		return (scanned,)


   @app.cell
   def _():
   		# Optional cleanup. To run uncomment and run cell
   		# print(f"Deleting table: {test_table}")
   		# catalog.drop_table(test_table)
   		# print("Table dropped.")
   		return


   if __name__ == "__main__":
   		app.run()
   ```


3. Replace the `CATALOG_URI`, `WAREHOUSE`, and `TOKEN` variables with your values from sections **1** and **2** respectively.
4. Launch the notebook editor in your browser:

   ```plaintext
   uv run marimo edit r2-data-catalog-tutorial.py
   ```

   Once your notebook connects to the catalog, the catalog along with its namespaces and tables will appear in the Datasources panel in marimo.

In the Python notebook above, you:

1. Connect to your catalog.
2. Create the `default` namespace.
3. Create a simple PyArrow table.
4. Create (or load) the `people` table in the `default` namespace.
5. Append sample data to the table.
6. Print the contents of the table.
7. (Optional) Drop the `people` table we created for this tutorial.

## Learn more

### [Managing catalogs](https://641a99ee.previews.developers.cloudflare.com/r2-data-catalog/manage-catalogs/)

Enable or disable R2 Data Catalog on your bucket, retrieve configuration details, and authenticate your Iceberg engine.

### [Connect to Iceberg engines](https://641a99ee.previews.developers.cloudflare.com/r2-data-catalog/config-examples/)

Find detailed setup instructions for Apache Spark and other common query engines.

Was this helpful?

YesNo

## On this page

[![](https://641a99ee.previews.developers.cloudflare.com/_astro/logo.te5VL_aD.svg)Docs](https://641a99ee.previews.developers.cloudflare.com/)

```json
{"@context":"https://schema.org","@type":"TechArticle","@id":"https://developers.cloudflare.com/r2-data-catalog/get-started/#page","headline":"Getting started · Cloudflare R2 Data Catalog docs","description":"Learn how to enable the R2 Data Catalog on your bucket, load sample data, and run your first query.","url":"https://developers.cloudflare.com/r2-data-catalog/get-started/","inLanguage":"en","image":"https://developers.cloudflare.com/og-docs.png","dateModified":"2026-09-03","publisher":{"@type":"Organization","name":"Cloudflare","description":"One platform for your apps, agents, and workforce. Build, secure, and scale without managing infrastructure","url":"https://www.cloudflare.com/","sameAs":["https://github.com/cloudflare","https://www.linkedin.com/company/cloudflare","https://x.com/cloudflare"],"logo":{"@type":"ImageObject","url":"https://developers.cloudflare.com/logo.svg"},"address":{"@type":"PostalAddress","streetAddress":"101 Townsend St","addressLocality":"San Francisco","addressRegion":"CA","postalCode":"94107","addressCountry":"US"},"contactPoint":[{"@type":"ContactPoint","contactType":"Customer Support","url":"https://support.cloudflare.com/","availableLanguage":["English"]},{"@type":"ContactPoint","contactType":"Sales","url":"https://www.cloudflare.com/contact/","availableLanguage":["English"]}]},"isPartOf":{"@type":"WebSite","@id":"https://developers.cloudflare.com/#website","name":"Cloudflare Docs","url":"https://developers.cloudflare.com/"}}
```
