Point-in-Time Recovery — recover a Lakebase branch and repoint an app

A standalone, repurposable script that recovers Lakebase data to an earlier point in time using branching, then repoints a running Databricks App at the recovered data. Because Lakebase branches are copy-on-write, “recovery” is a branch created as of a past timestamp plus a one-variable redeploy — no restore job, no backup file, and the live branch is left untouched so you can compare before cutting over.

Reach for it after a bad deploy, a runaway UPDATE/DELETE, or any data corruption: create a branch from just before the damage, verify it, and point the app there.

Features

Area What you get
Point-in-time branch Fork a recovery branch off any source branch as of an exact instant or N hours ago — copy-on-write, ready in seconds
App repoint Rewrite the deployed app’s app.yaml Lakebase env vars (LAKEBASE_HOST + LAKEBASE_ENDPOINT) and redeploy (SNAPSHOT) — same app, same code, only the branch changes
Access re-grant Idempotently (re)grant the app’s service principal on the new branch, so the repointed app doesn’t fail auth
Branch-only mode --no-repoint creates and reports the branch (host + endpoint) without touching any app, for manual inspection first
Repurposable Pure Python + the Databricks SDK — no notebook, no dbutils, no repo-specific setup; every workspace value is an env var or CLI flag

Architecture

   source branch (e.g. production)          recovery branch
   ─────────────────────────────            ────────────────
   corrupted "now"                          clean, as of T-1h
          │                                        ▲
          │  create_branch(source_branch_time=T)   │ copy-on-write
          └────────────────────────────────────────┘
                                 │
                                 ▼
        rewrite app.yaml (LAKEBASE_HOST + LAKEBASE_ENDPOINT)
                                 │
                                 ▼
        grant app service principal on the recovery branch
                                 │
                                 ▼
        redeploy app (SNAPSHOT) ──►  app serves clean data

The script talks to two Databricks APIs: Postgres (create the branch, list its endpoint, mint a credential to run the grants) and Apps (read the deployed source path, rewrite app.yaml in the workspace, redeploy).

How the recovery runs

point_in_time_recovery.py drives the whole flow:

  1. Create the recovery branch — create_branch off the source (production) branch with source_branch_time set to the recovery instant and no_expiry=True, so the branch is a copy-on-write copy of the schema and data as of that point and survives until you delete it.
  2. Resolve the endpoint — list the branch’s endpoints and take the host, so you know where the recovered data now lives.
  3. Rewrite the deployed app.yaml — export the target app’s app.yaml from the workspace, substitute the LAKEBASE_HOST and LAKEBASE_ENDPOINT env values to the recovery branch, and import it back.
  4. Re-grant the service principal — a branch has its own Postgres roles, so the app’s service principal is (re)granted SELECT/USAGE on the recovery branch (idempotently) or the repointed app fails auth.
  5. Redeploy (SNAPSHOT) — redeploy the app so it picks up the rewritten app.yaml and serves clean data from the recovery branch.

Run it

This is a reference script you run against your own workspace. Prerequisites:

  • A Lakebase project with a source branch that has point-in-time history (e.g. production) and a primary endpoint.
  • The databricks CLI authenticated (databricks auth login) or DATABRICKS_HOST + DATABRICKS_TOKEN exported, and uv.
  • To use the repoint step: a deployed Databricks App whose app.yaml sets LAKEBASE_HOST and LAKEBASE_ENDPOINT env vars.
cd developer_experience/point_in_time_recovery
uv sync

# Configure — copy .env.example to .env and fill it in, or export the vars:
export LAKEBASE_PROJECT=<your-lakebase-project-id>
export DATABRICKS_APP_NAME=<your-app-name>

# Recover to one hour ago and repoint the app:
uv run scripts/point_in_time_recovery.py --hours-back 1

# Or recover to an exact instant, branch only (inspect before repointing):
uv run scripts/point_in_time_recovery.py \
  --recovery-time 2026-08-26T14:30:00Z --no-repoint

--no-repoint (or omitting --app-name / DATABRICKS_APP_NAME) creates the branch and prints its host and endpoint so you can connect and verify the data before cutting the app over.

Configuration

Every value is an env var with a matching CLI flag (the flag wins). Run uv run scripts/point_in_time_recovery.py --help for the full list.

Env var / flag Purpose Default
LAKEBASE_PROJECT / --project Lakebase project id. Required. —
LAKEBASE_SOURCE_BRANCH / --source-branch Branch to recover from. production
LAKEBASE_ENDPOINT / --endpoint Endpoint name on the branch. primary
LAKEBASE_DATABASE / --database Postgres database. databricks_postgres
LAKEBASE_PG_SCHEMA / --schema Schema the app reads (used for the grants). public
DATABRICKS_APP_NAME / --app-name App to repoint. Unset ⇒ branch only. —
--hours-back Recover to N hours before now (UTC). 1
--recovery-time Recover to an explicit ISO-8601 instant (overrides --hours-back). —
--branch-id Recovery branch id to create. recovery-<UTC timestamp>
--no-repoint Only create the branch; never touch an app. off

Authentication uses the standard Databricks SDK resolution — a --profile in ~/.databrickscfg, DATABRICKS_CONFIG_PROFILE, or DATABRICKS_HOST + DATABRICKS_TOKEN.

Notes and caveats

  • The grants are read-only (SELECT/USAGE). If your app writes to Lakebase, widen the grants in grant_app_access() accordingly.
  • The recovery branch is created with no_expiry so it survives until you delete it. Branches are copy-on-write and cheap, but delete the ones you no longer need. To cut a repointed app back to production, rerun the app’s normal deploy (it re-renders app.yaml), or repoint it at production the same way.
  • The repoint edits the deployed app.yaml in place. A subsequent databricks bundle deploy of the app re-renders app.yaml from source and reverts the repoint — expected, since production is the steady state.