Skip to content

cloud: add tiered storage docs - #23328

Open
zhaoshangzi wants to merge 55 commits into
pingcap:release-8.5from
zhaoshangzi:ts_prip_v1
Open

zhaoshangzi wants to merge 55 commits into
pingcap:release-8.5from
zhaoshangzi:ts_prip_v1

Conversation

@zhaoshangzi

@zhaoshangzi zhaoshangzi commented Jul 22, 2026 •

Copy link
Copy Markdown
Collaborator

First-time contributors' checklist

What is changed, added or deleted? (Required)

Which TiDB version(s) do your changes apply to? (Required)

Tips for choosing the affected version(s):

By default, CHOOSE MASTER ONLY so your changes will be applied to the next TiDB major or minor releases. If your PR involves a product feature behavior change or a compatibility change, CHOOSE THE AFFECTED RELEASE BRANCH(ES) AND MASTER.

For details, see tips for choosing the affected versions.

  • master (the latest development version)
  • v9.0 (TiDB 9.0 versions)
  • v8.5 (TiDB 8.5 versions)
  • v8.1 (TiDB 8.1 versions)
  • v7.5 (TiDB 7.5 versions)
  • v7.1 (TiDB 7.1 versions)
  • v6.5 (TiDB 6.5 versions)

What is the related PR or file link(s)?

  • This PR is translated from:
  • Other reference link(s):

AI agent involvement

  • The changes in this PR were primarily made by an AI agent on behalf of the PR author.

Do your changes match any of the following descriptions?

  • Delete files
  • Change aliases
  • Need modification after applied to another branch
  • Might cause conflicts after applied to another branch

Summary by CodeRabbit

  • Documentation
    • Added a Tiered Storage section to the TiDB Cloud documentation table of contents.
    • Introduced new pages covering Tiered Storage concepts, detailed operations and configuration, limitations/impact, and monitoring considerations.
    • Added a Tiered Storage FAQ covering update/delete support, caching/cold-read behavior, TiFlash behavior, outage impact, replica storage differences, compatibility considerations, and cost impact.
  • Chores
    • Updated local ignore rules to exclude local settings.

@ti-chi-bot ti-chi-bot Bot added contribution This PR is from a community contributor. missing-translation-status This PR does not have translation status info. size/XXL Denotes a PR that changes 1000+ lines, ignoring generated files. labels Jul 22, 2026
@coderabbitai

coderabbitai Bot commented Jul 22, 2026 •

Copy link
Copy Markdown

Review Change Stack

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review
📝 Walkthrough

Walkthrough

Adds TiDB Cloud Essential Private Preview documentation for Tiered Storage, including concepts, architecture, configuration, operations, limitations, FAQ content, and table-of-contents navigation. It also ignores a local QodEr settings file.

Changes

Tiered Storage documentation

Layer / File(s) Summary
Concepts and architecture
tidb-cloud/tiered-storage-concepts.md
Defines IA storage behavior, usage criteria, TiDB–TiKV–Object Store architecture, caching, read amplification, write paths, and conversion characteristics.
Configuration and operations
tidb-cloud/tiered-storage-operations.md
Documents storage classes, table and partition DDL, inspection, monitoring, observability, rollout, optimization, and switch-back practices.
Limitations and FAQ
tidb-cloud/tiered-storage-limitations.md, tidb-cloud/tiered-storage-faq.md
Documents feature restrictions, throttling, tool impacts, recovery methods, cache behavior, and IA storage questions.
Documentation navigation
TOC-tidb-cloud-byoc.md
Adds Concepts, Operations, Limitations, and FAQ links under Tiered Storage.

Local tooling ignore

Layer / File(s) Summary
QodEr settings exclusion
.gitignore
Ignores .qoder/settings.local.json.

Estimated code review effort: 2 (Simple) | ~10 minutes

Suggested labels: type/enhancement

Suggested reviewers: connor1996, lilin90

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Description check ⚠️ Warning The template is present, but the required content sections are blank and related links/notes are missing. Add a brief change summary, confirm affected version(s), include related links, and fill AI involvement and other checklist items as applicable.
✅ Passed checks (4 passed)
Check name Status Explanation
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title accurately summarizes the main change: adding tiered storage documentation to cloud docs.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 14


ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 8c969826-2f18-4a5e-9ea9-0868fa38a3a4

📥 Commits

Reviewing files that changed from the base of the PR and between 7182794 and fd164f0.

📒 Files selected for processing (5)
  • TOC-tidb-cloud.md
  • tidb-cloud/tieredstorage_concepts.md
  • tidb-cloud/tieredstorage_faq.md
  • tidb-cloud/tieredstorage_limitations.md
  • tidb-cloud/tieredstorage_operations.md

Comment thread tidb-cloud/tieredstorage_concepts.md Outdated

## 1 What is Tiered Storage

Tiered Storage is a **table-level and partition-level storage tiering capability** on TiDB Cloud Essential, designed for infrequently accessed data. Users can set a table or partition to the IA (Infrequent Access) storage class. The system automatically stores the full data in remote object storage (S3/OSS, etc.), keeping only metadata and on-demand cached hot data segments locally.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win

Replace etc. with explicit examples.

Vale rejects the current wording. Use a precise list without changing the technical meaning.

Suggested replacement
-Tiered Storage is a **table-level and partition-level storage tiering capability** on TiDB Cloud Essential, designed for infrequently accessed data. Users can set a table or partition to the IA (Infrequent Access) storage class. The system automatically stores the full data in remote object storage (S3/OSS, etc.), keeping only metadata and on-demand cached hot data segments locally.
+Tiered Storage is a **table-level and partition-level storage tiering capability** on TiDB Cloud Essential, designed for infrequently accessed data. Users can set a table or partition to the IA (Infrequent Access) storage class. The system automatically stores the full data in remote object storage, such as S3 or OSS, keeping only metadata and on-demand cached hot data segments locally.

As per path instructions, this is an exact replacement for the contiguous changed line.

📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
Tiered Storage is a **table-level and partition-level storage tiering capability** on TiDB Cloud Essential, designed for infrequently accessed data. Users can set a table or partition to the IA (Infrequent Access) storage class. The system automatically stores the full data in remote object storage (S3/OSS, etc.), keeping only metadata and on-demand cached hot data segments locally.
Tiered Storage is a **table-level and partition-level storage tiering capability** on TiDB Cloud Essential, designed for infrequently accessed data. Users can set a table or partition to the IA (Infrequent Access) storage class. The system automatically stores the full data in remote object storage, such as S3 or OSS, keeping only metadata and on-demand cached hot data segments locally.
🧰 Tools
🪛 GitHub Check: vale

[failure] 12-12:
[vale] reported by reviewdog 🐶
[PingCAP.Latin] Use 'such as' instead of 'etc.'.

Raw Output:
{"message": "[PingCAP.Latin] Use 'such as' instead of 'etc.'.", "location": {"path": "tidb-cloud/tieredstorage_concepts.md", "range": {"start": {"line": 12, "column": 310}}}, "severity": "ERROR"}

Sources: Path instructions, Linters/SAST tools

Comment thread tidb-cloud/tieredstorage_concepts.md Outdated

Tiered Storage is a **table-level and partition-level storage tiering capability** on TiDB Cloud Essential, designed for infrequently accessed data. Users can set a table or partition to the IA (Infrequent Access) storage class. The system automatically stores the full data in remote object storage (S3/OSS, etc.), keeping only metadata and on-demand cached hot data segments locally.

**In a nutshell**: An IA table is still a regular table from the application layer — all query, transaction, backup, and recovery semantics remain unchanged. The difference lies in cost and performance: local disk usage drops significantly, but on a cold read (when the local cache is missing), data must be fetched from remote object storage, resulting in higher latency than Standard tables.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win

Remove spaces around the em dash.

Suggested replacement
-**In a nutshell**: An IA table is still a regular table from the application layer — all query, transaction, backup, and recovery semantics remain unchanged. The difference lies in cost and performance: local disk usage drops significantly, but on a cold read (when the local cache is missing), data must be fetched from remote object storage, resulting in higher latency than Standard tables.
+**In a nutshell**: An IA table is still a regular table from the application layer—all query, transaction, backup, and recovery semantics remain unchanged. The difference lies in cost and performance: local disk usage drops significantly, but on a cold read (when the local cache is missing), data must be fetched from remote object storage, resulting in higher latency than Standard tables.

As per path instructions, this is an exact replacement for the contiguous changed line.

📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
**In a nutshell**: An IA table is still a regular table from the application layer — all query, transaction, backup, and recovery semantics remain unchanged. The difference lies in cost and performance: local disk usage drops significantly, but on a cold read (when the local cache is missing), data must be fetched from remote object storage, resulting in higher latency than Standard tables.
**In a nutshell**: An IA table is still a regular table from the application layer—all query, transaction, backup, and recovery semantics remain unchanged. The difference lies in cost and performance: local disk usage drops significantly, but on a cold read (when the local cache is missing), data must be fetched from remote object storage, resulting in higher latency than Standard tables.
🧰 Tools
🪛 GitHub Check: vale

[failure] 14-14:
[vale] reported by reviewdog 🐶
[PingCAP.EmDash] Don't put a space before or after a dash.

Raw Output:
{"message": "[PingCAP.EmDash] Don't put a space before or after a dash.", "location": {"path": "tidb-cloud/tieredstorage_concepts.md", "range": {"start": {"line": 14, "column": 83}}}, "severity": "ERROR"}

Sources: Path instructions, Linters/SAST tools

Comment thread tidb-cloud/tiered-storage-concepts.md Outdated
Comment thread tidb-cloud/tiered-storage-concepts.md Outdated
- [ ] The data access frequency of the table/partition has been confirmed to be declining from a business perspective
- [ ] The table can be changed to a partitioned table, because cold/hot data separation is easier to manage with partitioned tables
- [ ] For regular table cold/hot separation, hot data accounts for less than 10% of the table
- [ ] Cold data access frequency is very low, e.g., query QPS does not exceed 10 concurrent (to avoid saturating object storage bandwidth)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🚀 Performance & Scalability | 🟠 Major | ⚡ Quick win

Separate QPS, concurrency, and throughput limits.

These lines currently treat “10 concurrent” as equivalent to a QPS or throughput limit. Define each dimension independently and explain the assumptions behind any conversion between them; otherwise users cannot determine whether their workload is within the documented safety limits.

  • tidb-cloud/tieredstorage_concepts.md#L60-L60: replace the QPS-based wording with an explicit concurrent-request limit.
  • tidb-cloud/tieredstorage_limitations.md#L30-L30: document the 1 GiB/s throughput limit separately from the maximum concurrent request count.
📍 Affects 2 files
  • tidb-cloud/tieredstorage_concepts.md#L60-L60 (this comment)
  • tidb-cloud/tieredstorage_limitations.md#L30-L30

Comment thread tidb-cloud/tieredstorage_faq.md Outdated
- **Standard layer**: Three replicas are each stored on the local disks of three TiKV nodes
- **IA layer**: Three replicas are each uploaded to object storage independently in IA format

Each replica on each TiKV node runs its own independent LSM-Tree, performing independent flush and compaction operations. When a table switches to IA storage class, all subsequently generated SST files are written to S3 in IA type. The three replicas produce their own independent SST files — three separate objects in S3, not shared.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Restrict IA SSTable writes to the documented L1+ path.

The concepts document explicitly says memtable and L0 remain local; only eligible L1+ files are opened in IA mode after flush or compaction.

Suggested replacement
-Each replica on each TiKV node runs its own independent LSM-Tree, performing independent flush and compaction operations. When a table switches to IA storage class, all subsequently generated SST files are written to S3 in IA type. The three replicas produce their own independent SST files — three separate objects in S3, not shared.
+Each replica on each TiKV node runs its own independent LSM-tree, performing independent flush and compaction operations. Memtable and L0 writes remain on the local hot path; after flush or compaction, eligible L1+ SST files are written to S3 in IA format. The three replicas produce their own independent SST files—three separate objects in S3, not shared.

As per path instructions, this is an exact replacement for the contiguous changed line.

📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
Each replica on each TiKV node runs its own independent LSM-Tree, performing independent flush and compaction operations. When a table switches to IA storage class, all subsequently generated SST files are written to S3 in IA type. The three replicas produce their own independent SST files — three separate objects in S3, not shared.
Each replica on each TiKV node runs its own independent LSM-tree, performing independent flush and compaction operations. Memtable and L0 writes remain on the local hot path; after flush or compaction, eligible L1+ SST files are written to S3 in IA format. The three replicas produce their own independent SST files—three separate objects in S3, not shared.
🧰 Tools
🪛 GitHub Check: vale

[failure] 43-43:
[vale] reported by reviewdog 🐶
[PingCAP.EmDash] Don't put a space before or after a dash.

Raw Output:
{"message": "[PingCAP.EmDash] Don't put a space before or after a dash.", "location": {"path": "tidb-cloud/tieredstorage_faq.md", "range": {"start": {"line": 43, "column": 291}}}, "severity": "ERROR"}

Source: Path instructions

Comment thread tidb-cloud/tieredstorage_faq.md Outdated

Each replica on each TiKV node runs its own independent LSM-Tree, performing independent flush and compaction operations. When a table switches to IA storage class, all subsequently generated SST files are written to S3 in IA type. The three replicas produce their own independent SST files — three separate objects in S3, not shared.

IA is a **storage format/location optimization**, not a **replica reduction mechanism**. It changes "how each node stores its own copy," not "how many copies exist." The Raft write and replication flow is identical to the Standard layer.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win

Replace the ambiguous “how many copies exist.”

Suggested replacement
-IA is a **storage format/location optimization**, not a **replica reduction mechanism**. It changes "how each node stores its own copy," not "how many copies exist." The Raft write and replication flow is identical to the Standard layer.
+IA is a **storage format/location optimization**, not a **replica reduction mechanism**. It changes how each node stores its own copy, not the number of copies. The Raft write and replication flow is identical to the Standard layer.

As per path instructions, this is an exact replacement for the contiguous changed line.

📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
IA is a **storage format/location optimization**, not a **replica reduction mechanism**. It changes "how each node stores its own copy," not "how many copies exist." The Raft write and replication flow is identical to the Standard layer.
IA is a **storage format/location optimization**, not a **replica reduction mechanism**. It changes how each node stores its own copy, not the number of copies. The Raft write and replication flow is identical to the Standard layer.
🧰 Tools
🪛 GitHub Check: vale

[warning] 45-45:
[vale] reported by reviewdog 🐶
[PingCAP.Ambiguous] Consider using a clearer word than 'many' because it may cause confusion.

Raw Output:
{"message": "[PingCAP.Ambiguous] Consider using a clearer word than 'many' because it may cause confusion.", "location": {"path": "tidb-cloud/tieredstorage_faq.md", "range": {"start": {"line": 45, "column": 147}}}, "severity": "INFO"}

Sources: Path instructions, Linters/SAST tools

Comment thread tidb-cloud/tiered-storage-limitations.md Outdated
Comment thread tidb-cloud/tiered-storage-operations.md Outdated
Comment thread tidb-cloud/tiered-storage-operations.md Outdated
@lilin90 lilin90 changed the title add tiered storage contents 1 cloud: add tiered storage docs Jul 22, 2026
@lilin90 lilin90 self-assigned this Jul 22, 2026
@lilin90 lilin90 added the translation/no-need No need to translate this PR. label Jul 22, 2026
@ti-chi-bot ti-chi-bot Bot removed the missing-translation-status This PR does not have translation status info. label Jul 22, 2026
@lilin90 lilin90 added area/tidb-cloud This PR relates to the area of TiDB Cloud. BYOC Related to TiDB Cloud BYOC docs. labels Jul 22, 2026
@ti-chi-bot

ti-chi-bot Bot commented Jul 22, 2026

Copy link
Copy Markdown

[APPROVALNOTIFIER] This PR is NOT APPROVED

This pull-request has been approved by:
Once this PR has been reviewed and has the lgtm label, please ask for approval from lilin90. For more information see the Code Review Process.
Please ensure that each of them provides their approval before proceeding.

The full list of commands accepted by this bot can be found here.

Details Needs approval from an approver in each of these files:

Approvers can indicate their approval by writing /approve in a comment
Approvers can cancel approval by writing /approve cancel in a comment

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 7

♻️ Duplicate comments (1)
tidb-cloud/tiered-storage-concepts.md (1)

66-66: ⚠️ Potential issue | 🟠 Major

Keep request concurrency and throughput as separate limits.

  • tidb-cloud/tiered-storage-concepts.md#L66-L66: replace the QPS wording with separate limits for no more than 10 concurrent cold-read requests and no more than 1 GiB/s aggregate throughput.
  • tidb-cloud/tiered-storage-limitations.md#L36-L36: split the combined 1 GiB/s (≤ 10 concurrent) value into separate throughput and concurrency rows.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: ac2090ff-f7bd-475c-a51f-7f1766f3ff3e

📥 Commits

Reviewing files that changed from the base of the PR and between fd164f0 and 2f8fe85.

📒 Files selected for processing (6)
  • .gitignore
  • TOC-tidb-cloud-byoc.md
  • tidb-cloud/tiered-storage-concepts.md
  • tidb-cloud/tiered-storage-faq.md
  • tidb-cloud/tiered-storage-limitations.md
  • tidb-cloud/tiered-storage-operations.md

Comment thread tidb-cloud/tiered-storage-overview.md Outdated
Comment thread tidb-cloud/tiered-storage-concepts.md Outdated
Comment thread tidb-cloud/tiered-storage-faq.md Outdated

## What happens to IA tables when the object store (S3) experiences an outage?

IA tables will be affected and become unavailable — since all data resides remotely, read requests must fetch from S3. Additionally, if S3 bandwidth is saturated, IA read/write performance will also be impacted.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟠 Major | ⚡ Quick win

Do not state that every IA table becomes unavailable during an object-store outage.

The concepts document says hot data can remain in the local IA cache. During an outage, uncached reads and operations requiring object storage can fail, but cached data is not necessarily unavailable.

Suggested replacement
-IA tables will be affected and become unavailable — since all data resides remotely, read requests must fetch from S3. Additionally, if S3 bandwidth is saturated, IA read/write performance will also be impacted.
+IA tables will be affected: reads for uncached data and operations requiring object storage may fail, while cached data may remain readable. If object-storage bandwidth is saturated, IA read/write performance will also be impacted.
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
IA tables will be affected and become unavailable — since all data resides remotely, read requests must fetch from S3. Additionally, if S3 bandwidth is saturated, IA read/write performance will also be impacted.
IA tables will be affected: reads for uncached data and operations requiring object storage may fail, while cached data may remain readable. If object-storage bandwidth is saturated, IA read/write performance will also be impacted.
🧰 Tools
🪛 GitHub Check: vale

[failure] 26-26:
[vale] reported by reviewdog 🐶
[PingCAP.EmDash] Don't put a space before or after a dash.

Raw Output:
{"message": "[PingCAP.EmDash] Don't put a space before or after a dash.", "location": {"path": "tidb-cloud/tiered-storage-faq.md", "range": {"start": {"line": 26, "column": 50}}}, "severity": "ERROR"}

Comment thread tidb-cloud/tiered-storage-faq.md Outdated
Comment thread tidb-cloud/tiered-storage-faq.md Outdated
Comment thread tidb-cloud/tiered-storage-limitations.md Outdated
Comment thread tidb-cloud/tiered-storage-operations.md Outdated

@Connor1996 Connor1996 left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why is the doc merged to release-8.5? Should be master

Comment thread tidb-cloud/tiered-storage-concepts.md Outdated
Comment thread tidb-cloud/tiered-storage-faq.md Outdated
Comment thread tidb-cloud/tiered-storage-operations.md Outdated
| Range Columns partitioned table | Supported | Must use `ENGINE_ATTRIBUTE` |
| List partitioned table | Supported | Must use `ENGINE_ATTRIBUTE` |
| List Columns partitioned table | Supported | Must use `ENGINE_ATTRIBUTE` |
| Hash partitioned table | **Not supported** | — |

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It has an ambiguity. Table-level IA is supported and applies to all partitions. TiDB rejects only partition-scoped storage-class definitions for HASH/KEY partitions.

Comment thread TOC-tidb-cloud-byoc.md Outdated
Comment thread .gitignore Outdated
lilin90 and others added 6 commits July 27, 2026 15:08
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
Comment thread tidb-cloud/tiered-storage-concepts.md Outdated
Comment thread tidb-cloud/tiered-storage-concepts.md Outdated
Comment thread tidb-cloud/tiered-storage-concepts.md Outdated
summary: Learn about Tiered Storage on TiDB Cloud BYOC/Premium/Essential, including its concepts, architecture, use cases, and read amplification.
---

# Tiered Storage Concepts

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
# Tiered Storage Concepts
# Tiered Storage Overview

@lilin90 lilin90 added the do-not-merge/hold Indicates that a PR should not merge because someone has issued a /hold command. label Aug 21, 2026
Comment thread tidb-cloud/tidb-cloud-billing.md Outdated
Comment thread tidb-cloud/tiered-storage-overview.md Outdated
Comment thread tidb-cloud/tiered-storage-overview.md

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Several moderate documentation inconsistencies and ambiguities remain unresolved.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 7 Medium severity

Open (7)
What changed in this PR

Adds private-preview Tiered Storage documentation for TiDB Cloud Premium and BYOC.

Changes:

  • Adds navigation entries and five Tiered Storage guides.
  • Documents configuration, monitoring, limitations, FAQs, and billing.
  • Adds IA cache-level billing guidance.
File Summary
TOC-tidb-cloud-premium.md Adds Tiered Storage navigation.
TOC-tidb-cloud-byoc.md Adds Tiered Storage navigation.
tidb-cloud/​tiered-storage-overview.md Documents concepts, architecture, and scenarios.
tidb-cloud/​tiered-storage-observability.md Documents monitoring and metrics.
tidb-cloud/​tiered-storage-limitations.md Documents limitations and operational risks.
tidb-cloud/​tiered-storage-guide.md Documents configuration and operations.
tidb-cloud/​tiered-storage-faq.md Answers common Tiered Storage questions.
tidb-cloud/​tidb-cloud-billing.md Documents IA cache billing effects.

💡 Configure MCP servers for context-aware, tailored reviews. Learn more in the docs.


## Can IA tables execute `UPDATE`/`DELETE`?

Yes. An `UPDATE` operation first loads the corresponding data from object storage into the IA cache, performs the modification, and writes a new SST file, the same flow as a regular `UPDATE`. Performance is affected by cold reads.
Comment on lines +32 to +36
Only one copy is stored on Amazon S3, and all three replicas share the same object.

In the cloud storage engine architecture, SST/blob data files have only one copy on object storage (S3/DFS) to begin with: files are uploaded once by flush/compaction, the S3 key contains no node/replica information, and the three Raft replicas reference the same file id through the Raft-replicated ChangeSet. The three-replica mechanism applies only to Raft logs, metadata, and each node's local cache, never to the data on object storage.

**Cost implications**: The storage volume on S3 is always about 1x the data size (it does not multiply with the replica count). What the IA tier saves is local disk usage on each node; data durability is guaranteed by the object storage itself, independent of the replica count.
Comment on lines +92 to +96
### Partitioned table DDL

Partitioned tables **do not support** the `STORAGE_CLASS` syntactic sugar and must use `ENGINE_ATTRIBUTE`.

Partition attributes support three selector types (cannot be mixed) plus a table-level default:
- New metrics:
- `Row-based IA Storage` — The storage space of data in the IA storage class
- `Row-based Standard Storage` — Total Standard table space
- Relationship: `Row-based Storage` = `Row-based IA Storage` + `Row-based Standard Storage`
Comment on lines +33 to +38
Since shared physical clusters have limited object storage bandwidth, IA cold storage access must comply with the following limits:

| Constraint dimension | Limit | Reason |
|-|-|-|
| Single SQL cold read throughput | ≤ 100 MiB/s | Prevents one query from consuming excessive bandwidth |
| Total concurrent cold read throughput | ≤ 1 GiB/s (≤ 10 concurrent) | Protects other tenants in the cluster |

- `COMPLETED_REPLICAS` increases and `LAST_UPDATE_TIME` keeps advancing: the conversion is progressing normally and the data volume is simply large.
- `DURATION` keeps growing but `COMPLETED_REPLICAS` does not increase for a long time: the conversion might be stuck because of a system exception, such as a TiKV rolling restart, temporarily insufficient resources, or short-term object storage unavailability.
- `TOTAL_REPLICAS`, `COMPLETED_REPLICAS`, `PROGRESS`, and `LAST_UPDATE_TIME` all stay `NULL`: no successful observation has been made yet. These columns are populated and cleared together, so `LAST_UPDATE_TIME` cannot tell whether observations are still being attempted. Check `DURATION` instead: it always increases while the conversion is being tracked. If these columns stay `NULL` while `DURATION` keeps growing, polling is still running but has not returned a valid observation.
- [ ] The data access frequency of the table/partition has been confirmed to be declining from a business perspective
- [ ] The table can be changed to a partitioned table, because cold/hot data separation is easier to manage with partitioned tables
- [ ] For regular table cold/hot separation, hot data accounts for less than 10% of the table
- [ ] Cold data access frequency is very low, e.g., query QPS does not exceed 10 concurrent (to avoid saturating object storage bandwidth)
@lilin90

lilin90 commented Sep 28, 2026

Copy link
Copy Markdown
Member

Note

As confirmed by @zhaoshangzi, we'll make the tiered storage docs public on docs.pingcap.com left navigation. So I updated the TOC files via dd40f09, f0c1159.

@lilin90

lilin90 commented Sep 28, 2026

Copy link
Copy Markdown
Member

As confirmed by @zhaoshangzi, we won't show tiered storage docs in the left navigation of docs.pingcap.com. So I moved related topics to ## _BUILD_ALLOWLIST to hide them via ed122d3.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/tidb-cloud This PR relates to the area of TiDB Cloud. BYOC Related to TiDB Cloud BYOC docs. contribution This PR is from a community contributor. do-not-merge/hold Indicates that a PR should not merge because someone has issued a /hold command. size/XXL Denotes a PR that changes 1000+ lines, ignoring generated files. translation/no-need No need to translate this PR.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants