Skip to content

cloud: add data pipeline documents to Lake - #23910

Open
ginkgoch wants to merge 14 commits into
pingcap:release-8.5from
ginkgoch:data-pipeline
Open

ginkgoch wants to merge 14 commits into
pingcap:release-8.5from
ginkgoch:data-pipeline

Conversation

@ginkgoch

@ginkgoch ginkgoch commented Sep 20, 2026 •

Copy link
Copy Markdown
Collaborator

First-time contributors' checklist

What is changed, added, or deleted? (Required)

Add data pipeline documents

Which TiDB version(s) do your changes apply to? (Required)

Tips for choosing the affected version(s):

By default, CHOOSE MASTER ONLY so your changes will be applied to the next TiDB major or minor releases. If your PR involves a product feature behavior change or a compatibility change, CHOOSE THE AFFECTED RELEASE BRANCH(ES) AND MASTER.

For details, see tips for choosing the affected versions.

  • master (the latest development version)
  • v9.0 (TiDB 9.0 versions)
  • v8.5 (TiDB 8.5 versions)
  • v8.1 (TiDB 8.1 versions)
  • v7.5 (TiDB 7.5 versions)
  • v7.1 (TiDB 7.1 versions)
  • v6.5 (TiDB 6.5 versions)

What is the related PR or file link(s)?

  • Related code change PR links (if applicable):
  • This PR is translated from:
  • Other reference link(s):

AI agent involvement

  • The changes in this PR were primarily made by an AI agent on behalf of the PR author.

Do your changes match any of the following descriptions?

  • Delete files
  • Change aliases
  • Need modification after applied to another branch
  • Might cause conflicts after applied to another branch

Summary by CodeRabbit

  • Documentation
    • Added documentation for TiDB Cloud Data Pipeline, including full snapshot export and continuous synchronization to TiDB Cloud Lake.
    • Added setup guides for Premium and Essential plans, including external-stage configuration for Amazon S3 and Alibaba Cloud OSS.
    • Added a support matrix covering DDL, DML, and data type compatibility.
    • Added FAQ guidance on external stages, event-driven ingestion, and billing.
    • Updated the Table of Contents with Data Pipeline documentation links.

@ti-chi-bot ti-chi-bot Bot added the contribution This PR is from a community contributor. label Sep 20, 2026
@ti-chi-bot

ti-chi-bot Bot commented Sep 20, 2026

Copy link
Copy Markdown

[APPROVALNOTIFIER] This PR is NOT APPROVED

This pull-request has been approved by:
Once this PR has been reviewed and has the lgtm label, please assign oreoxmt for approval. For more information see the Code Review Process.
Please ensure that each of them provides their approval before proceeding.

The full list of commands accepted by this bot can be found here.

Details Needs approval from an approver in each of these files:

Approvers can indicate their approval by writing /approve in a comment
Approvers can cancel approval by writing /approve cancel in a comment

@ti-chi-bot ti-chi-bot Bot added missing-translation-status This PR does not have translation status info. size/XXL Denotes a PR that changes 1000+ lines, ignoring generated files. labels Sep 20, 2026
@ginkgoch
ginkgoch requested a review from sdojjy September 20, 2026 10:04
@coderabbitai

coderabbitai Bot commented Sep 20, 2026 •

Copy link
Copy Markdown

Review Change StackReview Change Stack

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review
📝 Walkthrough

Walkthrough

The PR adds TiDB Cloud Data Pipeline documentation. It covers navigation, pipeline operation, external-stage configuration, support details, FAQs, and setup procedures for Premium and Essential instances.

Changes

Data Pipeline Documentation

Layer / File(s) Summary
Overview and reference documentation
TOC-tidb-cloud-premium.md, tidb-cloud/data-pipeline/data-pipeline-overview.md, tidb-cloud/data-pipeline/lake-data-pipeline-faq.md, tidb-cloud/data-pipeline/lake-data-pipeline-support-matrix.md
Adds navigation, pipeline operation details, availability information, FAQ content, and DDL, DML, and type support tables.
External-stage configuration
tidb-cloud/data-pipeline/lake-data-pipeline-configure-external-stage.md
Documents AWS S3 and Alibaba Cloud OSS setup, access policies, credentials, and optional SQS event-driven ingestion.
Premium pipeline setup
tidb-cloud/data-pipeline/setup-lake-data-pipeline-for-premium.md
Documents Premium pipeline creation, replication configuration, object selection, and edit, pause, resume, and delete behavior.
Essential pipeline setup
tidb-cloud/data-pipeline/setup-lake-data-pipeline-for-essential.md
Documents S3 access, CSV snapshot export, CDC changefeed creation through the Open API, and TiDB Cloud Lake integration setup.

Priority: ⬇️ Low

Estimated code review effort: 2 (Simple) | ~15 minutes

Change: Other

Merge Risk: 🟡 Moderate · up to 627cc

The setup guides can lead users to grant broader storage access than intended, and non-AWS Essential users may be unable to use the documented Role ARN path. These issues should be corrected before merge.

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check. Docstring coverage is scoped to functions touched by this diff. Analyzed 0 functions across 0…
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title clearly summarizes the main change: adding Data Pipeline documentation to TiDB Cloud Lake.
Description check ✅ Passed The description follows the template and identifies the documentation change. However, it does not select the affected TiDB version, and the change summary is brief.
✨ Finishing Touches 💡 1
🛠️ Fix failing CI checks 💡
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Create a new PR

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 14


ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository: pingcap/docs/.coderabbit.yaml

Review profile: ASSERTIVE

Plan: Advanced

Run ID: 41262998-272e-4187-8556-a224f9c9df77

📥 Commits

Reviewing files that changed from the base of the PR and between 884aa5a and 50a3d6f.

📒 Files selected for processing (7)
  • TOC-tidb-cloud-premium.md
  • tidb-cloud/data-pipeline/data-pipeline-overview.md
  • tidb-cloud/data-pipeline/lake-data-pipeline-configure-external-stage.md
  • tidb-cloud/data-pipeline/lake-data-pipeline-faq.md
  • tidb-cloud/data-pipeline/lake-data-pipeline-support-matrix.md
  • tidb-cloud/data-pipeline/setup-lake-data-pipeline-for-essential.md
  • tidb-cloud/data-pipeline/setup-lake-data-pipeline-for-premium.md

Included review availability: Your plan provides up to 4 included reviews per hour; 3 remain after this review.

Comment thread tidb-cloud/data-pipeline/data-pipeline-overview.md Outdated
Comment thread tidb-cloud/data-pipeline/lake-data-pipeline-configure-external-stage.md Outdated
Comment thread tidb-cloud/data-pipeline-lake-configure-external-stage-aws.md
Comment thread tidb-cloud/data-pipeline/lake-data-pipeline-faq.md Outdated
Comment thread tidb-cloud/data-pipeline/setup-lake-data-pipeline-for-essential.md Outdated
Comment thread tidb-cloud/data-pipeline-lake-setup-for-essential.md
Comment thread tidb-cloud/data-pipeline/setup-lake-data-pipeline-for-premium.md Outdated
Comment thread tidb-cloud/data-pipeline/setup-lake-data-pipeline-for-premium.md Outdated
Comment thread tidb-cloud/data-pipeline/setup-lake-data-pipeline-for-premium.md Outdated
Comment thread TOC-tidb-cloud-premium.md Outdated
Comment thread tidb-cloud/data-pipeline-lake-configure-external-stage.md Outdated
2. Read the warning and confirm the operation. Deleting a data pipeline:

- Immediately stops all data replication.
- Removes the TiDB Cloud Lake data source and integration task associated with the pipeline.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could we avoid stating this as unconditional? The delete workflow can finish while Lake cleanup is skipped when Lake refuses the request (for example, HTTP 403); the service explicitly leaves the Lake task/data source for manual cleanup in that case. Please say “attempts to remove” and mention that resources may remain if Lake refuses the request.


- The TiDB Cloud Lake warehouse must be in the **same region** as your {{{ .premium }}} instance.
- Only tables with a **primary key** can be replicated incrementally. If a table in the sync scope has no primary key, its incremental replication fails and an error is reported for that table.
- You can create up to 100 data pipelines per {{{ .premium }}} instance.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The service code does not enforce “100 data pipelines per Premium instance”. The only matching quota is bizChangefeedPerClusterQuota = 100; a full-only pipeline has no changefeed at all, while an incremental pipeline consumes one. Please either cite the separate product quota or describe the actual changefeed/cluster limit, otherwise this restriction is not aligned with the implementation.

- **Full Data + Incremental Data** (default): exports a full snapshot of the selected source data, and then continuously replicates row changes. This is the recommended mode for ongoing synchronization.
- **Full Data**: exports a one-time full snapshot of the selected source data only. No incremental data is replicated, and changes made on the source after the snapshot are ignored.

2. **Sync Interval**: the interval at which the data pipeline scans for new data. The default is `5 minutes`. Shorter intervals reduce data latency but increase the number of API calls to cloud storage.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is not the interval that the service directly uses to scan Lake. dataflow-service-ng treats the request as an end-to-end latency budget and splits it between TiCDC flush and Lake polling (producer 2–600s, consumer 5–300s); when SQS is configured, the consumer poll interval is set to zero. I also could not find a service default of 5 minutes. Please document the end-to-end semantics and source the default from the actual UI/API contract before stating it here.

## Restrictions

- The TiDB Cloud Lake warehouse must be in the **same region** as your {{{ .premium }}} instance.
- Only tables with a **primary key** can be replicated incrementally. If a table in the sync scope has no primary key, its incremental replication fails and an error is reported for that table.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please qualify the no-primary-key behavior. The pipeline does not reject these tables at creation: source validation returns them in the deny list, the snapshot path still counts/exports the selected set, and the generated TiCDC changefeed uses IGNORE_NOT_SUPPORT_TABLE. Whether a per-table error is shown is downstream behavior, so “incremental replication fails and an error is reported” needs evidence or more precise wording.


## Step 1: Create an OSS bucket

> 💡 If you already have an OSS bucket ready, skip this step — just make sure the bucket region matches the region of your TiDB Cloud instance.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The same-region requirement is not enforced for Alibaba OSS by dataflow-service-ng: its provider tests explicitly accept an OSS region unrelated to the cluster, and the region check is only applied to the AWS path. Please scope this sentence to AWS, or cite the separate Lake/product constraint if OSS is intentionally restricted there.


## DDL support summary

| DDL pattern | Status | Behavior / symptom |

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This support matrix cannot be derived from the Data Pipeline implementation as written. The service configures TiCDC CANAL_JSON/content-compatible output, IGNORE_NOT_SUPPORT_TABLE, and Lake AllowDelete; it has no DDL/type compatibility table or translation for the listed operations and types. Please cite the tested TiCDC/Lake version and source for these claims, or label this as a separately validated compatibility matrix rather than an implementation guarantee.

Comment thread tidb-cloud/data-pipeline-lake-configure-external-stage.md Outdated
| `DROP COLUMN` | ✅ | |
| `ADD INDEX` / `DROP INDEX` | ✅ | |
| `RENAME COLUMN` | ✅ | |
| `DROP TABLE` | N/A | Destination table is kept in Lake. |

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This behavior conflicts with the webapi/bend-cdc implementation/docs used by the manual Lake integration: bend-cdc/docs/tidb-column-ddl.md says DROP TABLE, TRUNCATE, and RENAME TABLE are unsupported and block subsequent consumption for the table, rather than being silently ignored while the destination keeps/splits data. Please reconcile this matrix with the deployed engine/version and state whether these events stop one table or are intentionally skipped.

- **Table Rules**: `*.*` to sync all exported tables.
- **Changefeed S3 Prefix**: `<prefix>/incremental/`.
- **Dumpling S3 Prefix**: `<prefix>/snapshot/`.
- **Poll Interval**: default or as needed. A shorter interval reduces data latency but increases Lake hosting cost.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please state the actual defaults here. In webapi/bend-cdc, the TiDB manager defaults PollInterval to 60 seconds and MergeInterval to 30 seconds when the API leaves them unset; these controls have different latency/cost effects. “Default or as needed” is too vague for a manual procedure and can make the documented behavior diverge from the deployed worker.


In the default mode, the Data Pipeline relies on two independent polling intervals — one on the Changefeed side (which periodically flushes incremental data to the external stage) and one on the Lake side (which periodically scans the external stage for new data). Because these two intervals do not coordinate, the effective end-to-end latency is higher than either interval alone.

In event-driven mode, the Changefeed still flushes data to the external stage on its configured interval, but each flush also triggers an **S3 event notification** to an SQS queue. Lake subscribes to this queue and loads new data as soon as the notification arrives, eliminating the additional latency caused by its own polling interval.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This description is not true for the Essential/manual path backed by webapi/bend-cdc: SQS is only a latency optimization that wakes an authoritative listing; the manager still runs scheduled polling and reports sqs+polling. Please scope this to the native dataflow-service-ng path (if that path really uses zero polling), or describe the two products separately so Essential users do not assume SQS replaces polling.


> **Note:**
>
> Only tables with a primary key can be replicated incrementally. Tables that lack a primary key are listed separately in the **Filter results** panel, and their incremental replication fails. Add a primary key to these tables before you create the data pipeline, or exclude them with filter rules such as `"!test.tbl1"`.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This now contradicts the restriction above: line 23 correctly says that no-primary-key tables are skipped for incremental replication, but this note still says their replication “fails”. The Data Pipeline creates the TiCDC changefeed with IGNORE_NOT_SUPPORT_TABLE, so please use the same “skipped/excluded from incremental replication” wording here; otherwise users receive two incompatible outcomes in one setup guide.

- **Role ARN**: the ARN from [Step 1](#step-1-create-the-role-with-export-cloudformation), or **Access Key ID** / **Secret Access Key** from [Use an Access Key](#use-an-access-key).
- **S3 Bucket Name**: the bucket name only (for example, `my-datapipeline-bucket`, not the full URI).
- **S3 Region**: the same region as your Essential instance.
4. **SQS Queue URL** is optional.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The manual flow exposes an optional queue URL but never gives Essential users the required queue setup: S3 bucket notification, queue policy, and ReceiveMessage/DeleteMessage/GetQueueAttributes permissions on the same Role ARN or access key. webapi/bend-cdc constructs its SQS consumer from those credentials and silently falls back to periodic listing when the queue is inaccessible, so a user can think event-driven ingestion is enabled when it is not. Please add the setup steps or link directly to the corresponding AWS SQS procedure.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1


ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository: pingcap/docs/.coderabbit.yaml

Review profile: ASSERTIVE

Plan: Advanced

Run ID: cf7bbcf8-dd5e-49f9-9717-9d9985c009b9

📥 Commits

Reviewing files that changed from the base of the PR and between fbb8be5 and 627cc94.

📒 Files selected for processing (7)
  • TOC-tidb-cloud-premium.md
  • tidb-cloud/data-pipeline/data-pipeline-overview.md
  • tidb-cloud/data-pipeline/lake-data-pipeline-configure-external-stage.md
  • tidb-cloud/data-pipeline/lake-data-pipeline-faq.md
  • tidb-cloud/data-pipeline/lake-data-pipeline-support-matrix.md
  • tidb-cloud/data-pipeline/setup-lake-data-pipeline-for-essential.md
  • tidb-cloud/data-pipeline/setup-lake-data-pipeline-for-premium.md

Included review availability: Your plan provides up to 4 included reviews per hour; 3 remain after this review.

Comment thread tidb-cloud/data-pipeline/setup-lake-data-pipeline-for-essential.md Outdated
@ti-chi-bot

ti-chi-bot Bot commented Sep 22, 2026

Copy link
Copy Markdown

@sdojjy: adding LGTM is restricted to approvers and reviewers in OWNERS files.

Details

In response to this:

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes-sigs/prow repository.

@lilin90 lilin90 added lake Related to TiDB Cloud Lake docs. translation/no-need No need to translate this PR. labels Sep 23, 2026
@ti-chi-bot ti-chi-bot Bot removed the missing-translation-status This PR does not have translation status info. label Sep 23, 2026
@lilin90
lilin90 requested a review from awxxxxxx September 23, 2026 07:34
@lilin90 lilin90 changed the title Add data pipeline documents. cloud: add data pipeline documents to Lake Sep 23, 2026
Comment thread tidb-cloud/data-pipeline-overview.md
Comment thread tidb-cloud/data-pipeline-lake-configure-external-stage-alibaba-cloud.md Outdated
Comment thread tidb-cloud/data-pipeline-lake-configure-external-stage-aws.md Outdated
Comment thread TOC-tidb-cloud-premium.md Outdated
@ti-chi-bot

ti-chi-bot Bot commented Sep 28, 2026

Copy link
Copy Markdown

@ginkgoch: The following test failed, say /retest to rerun all failed tests or /retest-required to rerun all mandatory failed tests:

Test name Commit Details Required Rerun command
pull-verify 262a2a0 link true /test pull-verify

Full PR test history. Your PR dashboard.

Details

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes-sigs/prow repository. I understand the commands that are listed here.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contribution This PR is from a community contributor. lake Related to TiDB Cloud Lake docs. size/XXL Denotes a PR that changes 1000+ lines, ignoring generated files. translation/no-need No need to translate this PR.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants