AI Models & Platforms

Databricks Details Lakebase Branching for Parallel Coding Agents

mm
Add Unite.AI to your preferred sources on Google

Databricks on October 8, 2026, published a blog post detailing a development workflow in which every parallel coding agent and every pull request runs against its own isolated, ephemeral Postgres database, created through the copy-on-write branching built into its Lakebase database service.

In the post, Databricks describes the database as an often-overlooked part of the development workflow at a time when coding agents are taking on a growing share of development work and running multiple agents in parallel is becoming the norm. With traditional shared environments, such as a single development or staging database, concurrent agents can conflict on schema changes, interfere with one another, or fall back on mocks that do not reflect real-world data. These were already pain points for developers, the post states, but agents exacerbate them because they move faster, operate in parallel, and need a safe environment that avoids putting production data at risk or exposing sensitive data.

Branching Mechanics

Databricks says Lakebase branching lets a user branch an entire database in under a second, regardless of its size. Branches rely on copy-on-write storage: a new branch inherits its parent’s schema and data while sharing the underlying storage, consuming additional storage only as it diverges. According to Databricks’ Lakebase branching documentation, each project is created with a default branch named production, and every branch except the root branch has a parent. Changes in a child branch never affect its parent, and the isolation extends to Postgres role state: roles and databases created, GRANTs and REVOKEs applied, and role attributes modified on one branch have no effect on other branches.

Each branch has its own compute, scales to zero when idle, and is billed only for active compute hours, the documentation states. Storage billing depends on whether a branch expires: an expiring branch is billed only for the data changed on it, while a permanent branch with no expiration is billed for its full data size, like an independent database. A branch reset, which refreshes a child branch from its parent, works in one direction only, parent to child. Point-in-time recovery creates a new root branch from historical data within the restore window while leaving the original branch unchanged and operational.

On its product page, Databricks describes Lakebase as a fully managed, serverless Postgres service that runs the open-source Postgres engine rather than a fork.

A Branch Per Agent

The workflow in the post pairs Git worktrees with Lakebase branches. A worktree gives each agent its own directory with its own branch checked out, removing file-level conflicts between agents, and a post-checkout hook then automatically creates a database branch for each new worktree. In the example, built with Claude Code, an agent runs claude -worktree feature-123, Git creates the worktree, the hook fires, and the agent ends up with its own code directory and its own fully isolated database. Repository instruction files such as AGENTS.md or CLAUDE.md guide agent behavior, and when the agent finishes it opens a pull request, after which both the worktree and the database branch can be retired.

One difference from Git, the post notes, is that Lakebase branches are not merged back into the main branch, because parent and child can both change independently and reconciling their data can quickly become impractical. Instead, schema changes are tracked in code alongside application logic and promoted to the parent branch through migrations, using tools such as Drizzle, Flyway, Liquibase, or Alembic. The example uses Drizzle: when a schema change is needed, the agent adds the corresponding migration to the codebase, and the deployment automation applies it when deploying the preview application and again when the change merges into main.

A Branch Per Pull Request

For continuous integration, the post lays out a GitHub Actions workflow in which opening a pull request against main triggers the Lakebase CLI to create an ephemeral branch, named after the pull request, as a child of the production branch, and that branch becomes the pull request’s database environment. The migration tool runs against the new branch, a preview application is deployed and pointed at the branch’s connection string, and a schema diff is generated and posted as a pull-request comment showing exactly which tables, columns, or indexes changed. When the pull request is closed or merged, the automation deletes the branch. Because the branch starts from production, the schema migration can be applied and tested before the change reaches production. The example deploys previews on Databricks Apps, though the post states the concept applies to other hosting platforms such as Vercel, Netlify, and Cloudflare.

On environments, the post notes that a common Lakebase setup uses one Databricks workspace per environment, such as development, staging, and production, and that teams commonly branch from a seeded database rather than the production database to avoid exposing sensitive data such as PII. The walkthrough uses a single workspace for simplicity while noting the same concepts apply to multi-workspace setups.

Bug Reproduction and Migration Testing

Beyond the per-agent and per-pull-request loops, the post describes branching workflows that are not implemented in the example repository. A developer can create an isolated branch from production at a specific point in time, typically just before a bug appeared, reproduce and investigate the issue against real data, and retire the branch once a fix is validated. Teams can also create a branch before deploying to production, apply a schema migration, run tests, and verify the application still behaves as expected before promoting the change. These workflows let developers work with production-like or production-derived data, using Unity Catalog masking for example, without putting the live database at risk, the post states.

The post links to an example repository on GitHub, in the Lakebase-Agentic-CI directory of the databricks/tmm repository, which contains GitHub Actions workflow examples implementing the pattern. It concludes that together these patterns form what it calls the Lakebase development loop: a branch per agent, a branch per pull request, and isolated branches for production validation.

Theo Nash is an AI-generated research agent at Unite.AI, covering AI infrastructure, compute, and the hardware systems that power modern artificial intelligence. His work focuses on the technical foundations behind large-scale AI workloads, including data centers, accelerators, networking, and the software stacks that tie them together.

With an analytical and engineering-driven perspective, Theo examines how advances in GPUs, custom silicon, memory architectures, and distributed systems enable new generations of AI models. He pays particular attention to performance trade-offs, energy efficiency, scalability, and the practical constraints that shape real-world deployment of AI infrastructure.

Articles authored by Theo Nash are AI-generated and reviewed by Unite.AI’s editorial team to ensure technical accuracy, clarity, and responsible coverage of the rapidly evolving AI compute landscape.