AI Lake

Beta Feature

AI Lake and its public interface for Apache Iceberg data ingestion are currently in beta. Their behavior, supported clients, and configuration may change in future releases.

AI Lake is GoodData’s managed data infrastructure purpose-built for analytics driven by AI. It provides managed storage and compute for the data that powers AI-driven Business Intelligence and Agentic Analytics experiences in GoodData Cloud.

Eligible customers with the AI Lake entitlement can use the public Apache Iceberg interface to ingest data in the open Iceberg table format. The ingested tables can be connected to GoodData Cloud through an AI Lake data source, added to the semantic layer, and used across metrics, visualizations, dashboards, and AI-driven analytics experiences.

Customers who manage their own data pipelines can write to AI Lake from their infrastructure with standard Iceberg-compatible tools. Customers who prefer a managed implementation can work with GoodData Professional Services.

Data Lakes and AI Lake

A data lake is a storage environment for large volumes of structured, semi-structured, and unstructured data. It provides flexible and scalable storage that can be used by different data processing and analytics tools.

AI Lake builds on this concept by combining open-format storage with managed analytical compute and GoodData’s semantic layer. It stores tables in the Apache Iceberg format and serves analytical queries directly over the stored data.

In GoodData product terminology:

  • AI Lake is the managed storage and compute service for analytics.
  • AI Lake entitlement determines whether an organization is eligible to use AI Lake.
  • AI Lake data source connects tables stored in AI Lake to a GoodData Cloud workspace.

Why AI Lake Uses Apache Iceberg

Apache Iceberg is an open table format designed for reliable analytics on data stored in cloud object storage. It provides the following benefits:

  • Multi-engine interoperability: Iceberg tables can be accessed by compatible processing engines and data engineering tools. AI Lake currently uses StarRocks as its high-performance compute engine for analytical queries.
  • Separation of storage and compute: Data is stored in open Parquet files in cloud object storage, while analytical compute is provided separately.
  • ACID transactions: Iceberg supports safe concurrent reads and writes without exposing partial table updates.
  • Schema and partition evolution: Supported schema and partition changes can be applied through table metadata without rewriting existing data files.
  • Snapshots and time travel: Iceberg keeps table snapshots that can support point-in-time reads and rollback workflows.

The public beta does not yet document every Iceberg operation as a stable supported contract. See Current Limitations.

How AI Lake Works

AI Lake exposes an Apache Iceberg REST Catalog endpoint. Your Iceberg client uses this endpoint for catalog operations such as creating, loading, and updating tables.

A typical workflow is:

  1. GoodData provisions AI Lake storage for an eligible organization and provides the storage ID.
  2. You create an AI Lake database instance and choose its database ID.
  3. AI Lake automatically creates:
    • an Apache Iceberg namespace with the same ID as the database instance
    • a corresponding AI Lake data source in GoodData Cloud
  4. You obtain a GoodData API token for the organization.
  5. You configure an Apache Iceberg client with the catalog endpoint and token.
  6. Your pipeline creates tables in the existing namespace and writes data in the Apache Iceberg format.
  7. You add the tables from the automatically created AI Lake data source to the logical data model and use them in the semantic layer.
  8. GoodData Cloud workspaces query the data through AI Lake.

The catalog endpoint has the following format:

https://<ORGANIZATION_HOST>/api/v1/ailake/database/<DATABASE_ID>/catalog

Before You Start

You need:

  • a GoodData Cloud organization with the AI Lake entitlement
  • an AI Lake storage ID provided by GoodData
  • a GoodData API token
  • the cloud region provided during onboarding
  • a supported Apache Iceberg client

GoodData provisions the AI Lake storage. You then use the public API to create database instances in that storage.

The current beta has been validated with:

  • PyIceberg
  • Apache Spark with the Iceberg Spark runtime

Other clients that support the Iceberg REST Catalog protocol may work, but they have not been validated for the current beta.

Create an AI Lake Database and Data Source

Before you write Iceberg tables, create an AI Lake database instance.

Choose the value of DATABASE_ID. When the database instance is created, AI Lake automatically creates:

  • an Iceberg namespace with the same ID
  • a corresponding AI Lake data source in GoodData Cloud

Do not create the namespace separately with your Iceberg client. You also do not need to create the AI Lake data source separately.

Set the required variables:

export HOSTNAME="<ORGANIZATION_HOST>"
export TOKEN="<GOODDATA_API_TOKEN>"
export DATABASE_ID="<DATABASE_ID>"
export STORAGE_ID="<STORAGE_ID>"

Create the database instance:

curl "https://${HOSTNAME}/api/v1/ailake/database/instances" \
  -X POST \
  -H "Accept: application/json, text/plain, */*" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer ${TOKEN}" \
  --data-raw "{
    \"name\": \"${DATABASE_ID}\",
    \"storageIds\": [
      \"${STORAGE_ID}\"
    ]
  }"

After the database is created:

  • use DATABASE_ID in the Iceberg REST Catalog URL
  • use the same DATABASE_ID as the Iceberg namespace
  • create tables directly in this namespace
  • use the corresponding AI Lake data source to add the tables to the logical data model

For example, if DATABASE_ID is analytics, create the table as analytics.events. Do not run create_namespace("analytics") or CREATE NAMESPACE analytics, because the namespace already exists.

Write Data with PyIceberg

Install PyIceberg with the required AWS and PyArrow dependencies:

pip install "pyiceberg[pyarrow,s3fs]"

Configure the REST catalog and write sample data:

from pyiceberg.catalog import load_catalog
import pyarrow as pa

database_id = "<DATABASE_ID>"

catalog = load_catalog(
    "gooddata",
    **{
        "type": "rest",
        "uri": f"https://<ORGANIZATION_HOST>/api/v1/ailake/database/{database_id}/catalog",
        "token": "<API_TOKEN>",
        "rest.sigv4-enabled": "false",
        "s3.region": "<AWS_REGION>",
    },
)

catalog.create_table(
    f"{database_id}.events",
    schema=pa.schema(
        [
            ("event_id", pa.int64()),
            ("event_name", pa.string()),
        ]
    ),
)

# Reload the table before the first write.
table = catalog.load_table(f"{database_id}.events")

table.append(
    pa.table(
        {
            "event_id": [1, 2, 3],
            "event_name": ["created", "updated", "completed"],
        }
    )
)

Important Notice

In the current beta, reload a newly created table before the first write. The credentials required for writing data are provided when the table is loaded, not when it is created.

Write Data with Apache Spark

Configure a Spark Iceberg catalog:

spark.sql.catalog.gooddata=org.apache.iceberg.spark.SparkCatalog
spark.sql.catalog.gooddata.type=rest
spark.sql.catalog.gooddata.uri=https://<ORGANIZATION_HOST>/api/v1/ailake/database/<DATABASE_ID>/catalog
spark.sql.catalog.gooddata.token=<API_TOKEN>
spark.sql.catalog.gooddata.io-impl=org.apache.iceberg.aws.s3.S3FileIO
spark.sql.catalog.gooddata.client.region=<AWS_REGION>
spark.sql.catalog.gooddata.cache-enabled=false
spark.sql.catalog.gooddata.rest-metrics-reporting-enabled=false

Create a table in the existing namespace and insert sample data:

CREATE TABLE gooddata.<DATABASE_ID>.events (
    event_id BIGINT,
    event_name STRING
)
USING iceberg;

INSERT INTO gooddata.<DATABASE_ID>.events
VALUES
    (1, 'created'),
    (2, 'updated'),
    (3, 'completed');

Keep the Spark catalog cache disabled in the current beta so Spark reloads the table before writing to it.

Use the AI Lake Data Source

When you create an AI Lake database instance, GoodData automatically creates the corresponding AI Lake data source.

AI Lake data sources are available only to organizations with the AI Lake entitlement. You cannot create an AI Lake data source manually in the GoodData Cloud UI.

After you write tables to the database through the Iceberg REST Catalog interface, use the automatically created data source to:

  • add the Iceberg tables to the logical data model
  • define facts, attributes, relationships, and datasets
  • create reusable metrics in the semantic layer
  • use the modeled data in visualizations, dashboards, AI-driven Business Intelligence, and Agentic Analytics

The database instance, Iceberg namespace, and AI Lake data source are created as part of the same provisioning flow. The database ID determines the namespace that you use when creating Iceberg tables.

DATETIME and Time Zone Conversion

AI Lake stores DATETIME values without a time zone. In the AI Lake data source configuration you choose how those values are interpreted:

InterpretationLDM data typeTime zone conversion
Local time zoneTIMESTAMPNone. Values are used as stored.
UTCTIMESTAMP_TZValues are treated as UTC and converted to the effective time zone. That time zone comes from the organization, workspace, or user setting, or from the execution, for example a dashboard time zone.

When you add tables to the logical data model, new date dataset references for DATETIME columns get the LDM data type that matches this data source setting.

You can override the data type on an individual reference through the logical model API or in YAML. Set data_type to TIMESTAMP or TIMESTAMP_TZ on the date dataset reference. The web-based modeler cannot change this type, but it keeps an override you make through the API or YAML.

Use AI Lake in Existing Pipelines

You can integrate the Iceberg REST Catalog endpoint into orchestration and transformation workflows that run on your infrastructure. For example, you can use Airflow, Dagster, Spark, or custom code to prepare data and write it to AI Lake.

When you use the public interface directly, you manage the pipeline, scheduling, and compute on your infrastructure. GoodData Professional Services can also help with managed ingestion implementations for customers who do not want to operate the pipeline themselves.

Roll Back an Iceberg Table

You can restore an AI Lake table to a previous retained Apache Iceberg snapshot after a transformation or load writes incorrect data.

Every write creates a new Iceberg snapshot, and the table’s main branch points to the current snapshot. A rollback moves that pointer to an earlier snapshot. The operation changes table metadata only. It does not copy or delete table data, and you can reverse it while the relevant snapshots are still retained.

Rollback uses the existing AI Lake Iceberg REST Catalog. There is no separate rollback API.

Before You Roll Back

You need:

  • an organization with the AI Lake entitlement
  • a GoodData API token for a user with organization-manage permission
  • the AI Lake database ID or name
  • the namespace and table name

Set the variables used in the examples:

GD_HOST="https://<ORGANIZATION_HOST>"
GD_TOKEN="<API_TOKEN>"
DB="<DATABASE_ID>"
NS="<NAMESPACE>"
TBL="<TABLE>"
CATALOG="$GD_HOST/api/v1/ailake/database/$DB/catalog/v1/catalogs/ailake"
AUTH=(-H "Authorization: Bearer $GD_TOKEN")

The namespace must match the database namespace. Using another namespace returns 404.

Inspect Available Snapshots

Load the table metadata:

curl -s "${AUTH[@]}" "$CATALOG/namespaces/$NS/tables/$TBL" \
  | jq '{
      current: .metadata."current-snapshot-id",
      snapshots: [ .metadata.snapshots[] | {
          id: ."snapshot-id",
          at: (.["timestamp-ms"] / 1000 | todate),
          op: .summary.operation,
          records: .summary."total-records",
          files: .summary."total-data-files"
      } ],
      log: [ .metadata."snapshot-log"[] | {
          at: (.["timestamp-ms"] / 1000 | todate),
          id: ."snapshot-id"
      } ]
    }'

The snapshots array contains the snapshots that are still retained. The snapshot-log array shows the history of the main branch in commit order.

Important Notice

AI Lake currently retains snapshots for up to 120 hours. This retention period cannot be changed. If the state you want is no longer listed in snapshots, you cannot roll back to it.

To list the tables in the namespace:

curl -s "${AUTH[@]}" "$CATALOG/namespaces/$NS/tables" | jq

Choose a Recovery Snapshot

Choose the last snapshot before the incorrect write. Compare the snapshot timestamp, operation, and record count with the load or transformation you want to undo.

A snapshot with operation replace is commonly created by automatic file compaction. It is a valid rollback target, but it represents the table state at the time the compaction ran.

You can inspect a retained snapshot with PyIceberg before changing the current table state:

from pyiceberg.catalog.rest import RestCatalog

catalog = RestCatalog(
    "ailake",
    uri=f"{GD_HOST}/api/v1/ailake/database/{DB}/catalog",
    token=GD_TOKEN,
)

table = catalog.load_table((NS, TBL))
table.scan(snapshot_id=<SNAPSHOT_ID>).to_arrow().num_rows

Roll Back the Table

Use an Iceberg updateTable commit to point main to the chosen snapshot:

CURRENT=<CURRENT_SNAPSHOT_ID>
TARGET=<TARGET_SNAPSHOT_ID>

curl -s -X POST "${AUTH[@]}" -H 'Content-Type: application/json' \
  "$CATALOG/namespaces/$NS/tables/$TBL" \
  -d "{
    \"requirements\": [
      {\"type\": \"assert-ref-snapshot-id\", \"ref\": \"main\", \"snapshot-id\": $CURRENT}
    ],
    \"updates\": [
      {\"action\": \"set-snapshot-ref\", \"ref-name\": \"main\", \"type\": \"branch\",
       \"snapshot-id\": $TARGET}
    ]
  }"

The assert-ref-snapshot-id requirement protects against concurrent writes. If another commit changes the table after you inspect it, the catalog rejects the rollback with 409 instead of discarding the newer write.

With PyIceberg, use:

table.manage_snapshots().set_current_snapshot(snapshot_id=<TARGET_SNAPSHOT_ID>).commit()

Use set_current_snapshot, not rollback_to_snapshot. rollback_to_snapshot requires the target snapshot to be an ancestor of the current state, which can prevent a later undo or reject a valid retained snapshot when the parent chain is incomplete.

Verify the Rollback

Check the current snapshot after the update:

curl -s "${AUTH[@]}" "$CATALOG/namespaces/$NS/tables/$TBL" \
  | jq '.metadata."current-snapshot-id"'

The returned value should match the target snapshot. Then query the table again and verify that the data matches the recovery point.

The rollback is recorded in snapshot-log.

Important Notice

SQL query results can briefly show the previous state because the query engine caches table metadata. Re-run the query if the result appears stale immediately after a rollback.

Undo a Rollback

A rollback is reversible while both snapshots are retained. Repeat the rollback request with the snapshot IDs swapped: use the rollback target as the current snapshot requirement and the snapshot you rolled away from as the new target.

Rollback Errors

ResponseMeaningWhat to Do
409 CommitFailedExceptionThe table changed after you inspected it.Inspect the table again and rebuild the request with the new current snapshot ID. Do not remove assert-ref-snapshot-id.
404 NoSuchNamespaceExceptionThe namespace does not match the database namespace.Use the database’s namespace.
404 naming the snapshot IDThe target snapshot is no longer retained.Choose a retained snapshot. If the required state has expired, it cannot be recovered by rollback.
403 ForbiddenExceptionThe organization is not entitled to AI Lake, or the caller lacks organization-manage permission.Use an entitled organization and a token for a user with organization-manage permission.
401The token is missing or invalid.Check the token and the organization hostname.
502 ServiceUnavailableExceptionThe upstream catalog returned an error.Retry. If the error persists, contact GoodData Support with the timestamp.
400 BadRequestException naming tags, branches, or history.expire.*The request would create an unsupported table state that interferes with automatic snapshot maintenance.Do not create additional branches or tags, and do not change the history.expire.* properties.

Avoid Unsupported Table States

Do not create Iceberg tags or branches other than main to preserve a recovery point. Use the snapshot ID instead.

Do not set these table properties:

  • history.expire.max-snapshot-age-ms
  • history.expire.min-snapshots-to-keep

These settings do not extend the AI Lake snapshot retention window. They interfere with automatic snapshot expiry and can cause old snapshots to accumulate.

If a table already contains tags, non-main branches, or these history.expire.* properties, remove them and perform a normal write to the table. Automatic snapshot maintenance resumes on a later maintenance pass. Allow several hours. If the snapshot count has not decreased after a day, contact GoodData Support.

Current Beta Scope

The current beta focuses on the public interface for ingestion in the Apache Iceberg format. The documented and validated workflow includes:

  • authenticating with a GoodData API token
  • creating an AI Lake database instance in storage provisioned by GoodData
  • using the automatically created namespace whose ID matches the database ID
  • using the automatically created AI Lake data source
  • connecting to the public Iceberg REST Catalog endpoint
  • creating and loading Iceberg tables
  • appending data
  • inspecting retained Iceberg snapshots and rolling a table back to a retained snapshot
  • integrating the tables into the semantic layer
  • querying the modeled data through GoodData Cloud

Overwrite, delete, and schema evolution behavior is not yet documented as a stable beta contract. Test these operations thoroughly before using them in your pipelines.

Current Limitations

The following limitations apply during the beta:

  • AI Lake is available only as a GoodData-hosted capability in GoodData Cloud.
  • Your organization must have the AI Lake entitlement.
  • GoodData provisions the AI Lake storage and provides its storage ID. There is currently no public API for creating the storage itself.
  • You create AI Lake database instances through the public API.
  • AI Lake data sources are created automatically with database instances and cannot be created manually in the GoodData Cloud UI.
  • Only GoodData-managed storage and catalogs are supported.
  • Each AI Lake database instance has an Iceberg namespace with the same ID. Do not create that namespace separately.
  • PyIceberg and Spark are the only clients currently validated.
  • Streaming and sub-minute change-data-capture ingestion are not supported.
  • The public ingestion interface does not itself provide orchestration, scheduling, or source-specific connectors.
  • When using the public interface directly, you are responsible for the infrastructure and operation of your pipeline.
  • Reload a table before its first write. For Spark, keep the catalog cache disabled.
  • Iceberg snapshots are retained for up to 120 hours. The retention period cannot currently be changed.
  • Iceberg tags, branches other than main, and the history.expire.max-snapshot-age-ms and history.expire.min-snapshots-to-keep table properties are not supported for managing rollback retention.
  • The supported schema evolution, overwrite, and delete contract may change.