Manage Topics
Topics provide a way to organize events in a data streaming platform.
Create a topic
Creating a topic can be as simple as specifying a name for your topic on the command line. For example, to create a topic named xyz, run:
rpk topic create xyz
This command creates a topic named xyz with one partition and three replicas, because these are the default values set in the cluster configuration file. Replicas are copies of partitions that are distributed across different brokers, so if one broker goes down, other brokers still have a copy of the data.
Redpanda Cloud supports 40,000 topics per cluster.
Choose the number of partitions
A partition acts as a log file where topic data is written. Dividing topics into partitions allows producers to write messages in parallel and consumers to read messages in parallel. The higher the number of partitions, the greater the throughput.
| As a general rule, select a number of partitions that corresponds to the maximum number of consumers in any consumer group that will consume the data. |
For example, suppose you plan to create a consumer group with 10 consumers. To create topic xyz with 10 partitions, run:
rpk topic create xyz -p 10
Update topic configurations
After you create a topic, you can update the topic property settings for all new data written to it. For example, you can add partitions or change the cleanup policy.
Add partitions
You can assign a certain number of partitions when you create a topic, and add partitions later. For example, suppose you add brokers to your cluster, and you want to take advantage of the additional processing power. To increase the number of partitions for existing topics, run:
rpk topic add-partitions [TOPICS...] --num [#]
Note that --num <#> is the number of partitions to add, not the total number of partitions.
| If a topic already has messages and you add partitions, the existing messages won’t be redistributed to the new partitions. If you require messages to be redistributed, then you must create a new topic with the new partition count, then stream the messages from the old topic to the new topic so they are appropriately distributed according to the new partition hashing. |
Reduce the number of partitions
You cannot reduce the number of partitions on an existing topic. The Kafka API does not support it: a record’s partition is chosen when the record is produced, and records that have already been written stay in the partition where they landed, so removing a partition would orphan its data. Redpanda assigns a record that has a key to a partition by hashing the key, and leaves a record without a key to the producer’s partitioner, which usually spreads such records across all available partitions.
To move a topic’s data to fewer partitions, copy it to a new topic and switch your clients over. Before you start, check that your applications can tolerate the following:
-
Duplicates: Redpanda Connect delivers records at least once, so a restart or a retry during the copy can write the same record to the new topic twice. A lag of zero shows only how far the copy has committed, not that the new topic is free of duplicates. Either make your consumers idempotent, or deduplicate on a record ID after the copy.
-
Ordering: records keep their keys, so all records for a key still land on one partition and keep their order relative to each other, but the global order of records across partitions is not preserved.
-
Retention: the copy starts at the oldest record that is still retained. Records that retention or compaction has already removed cannot be copied, and both keep running during the copy, so complete the copy well within the topic’s retention period.
-
Consumer offsets: consumer group offsets are stored per topic, so the offsets your consumers committed on the original topic do not carry over. Each consumer starts from the beginning of the new topic and replays what the copy wrote, unless you set its offsets explicitly with
rpk group seek. -
Write downtime: producers must stop writing to the original topic before you switch clients over, so plan a window in which the topic accepts no writes.
The following procedure reduces a topic named orders from three partitions to one.
-
Check the configuration of the original topic so that you can recreate it. Note the replication factor, and every row whose
SOURCEisDYNAMIC_TOPIC_CONFIG, which is an override you must set on the new topic:rpk topic describe ordersExample output (abbreviated)
SUMMARY ======= NAME orders PARTITIONS 3 REPLICAS 1 CONFIGS ======= KEY VALUE SOURCE cleanup.policy delete DEFAULT_CONFIG retention.bytes -1 DEFAULT_CONFIG retention.local.target.ms 86400000 DEFAULT_CONFIG retention.ms 604800000 DYNAMIC_TOPIC_CONFIG segment.bytes 134217728 DEFAULT_CONFIGHere, only
retention.msis an override. If the topic has Tiered Storage settings, a custom cleanup policy, or other overrides, carry all of them over: a new topic created without them silently falls back to the cluster defaults. -
Create the new topic with the target number of partitions, the replication factor of the original topic, and each override from the previous step:
rpk topic create orders-reduced --partitions 1 --replicas 1 --topic-config retention.ms=604800000Example output
TOPIC STATUS orders-reduced OK -
Copy the data with a Redpanda Connect pipeline. This configuration reads all records that are still available in the original topic and preserves record keys, so records for the same key land on the same partition of the new topic:
reduce-partitions.yamlinput: redpanda: seed_brokers: ["<broker-address>"] topics: ["orders"] consumer_group: orders-to-orders-reduced start_offset: earliest output: redpanda: seed_brokers: ["<broker-address>"] topic: orders-reduced key: ${! @kafka_key }Give the consumer group a name that is unique to this copy, such as
<original-topic>-to-<new-topic>.start_offset: earliestapplies only when the group has no committed offset, so a group name that has been used before resumes from where it left off and skips records.rpk connect run reduce-partitions.yamlExample output
level=info msg="Launching a Redpanda Connect instance, use CTRL+C to close" level=info msg="Output type redpanda is now active" level=info msg="Input type redpanda is now active" -
Stop the producers that write to the original topic. Leave the pipeline running so that it copies the last records they wrote.
-
Wait for the copy to drain. It is complete when the consumer group reports a
LAGof0for every partition of the original topic:rpk group describe orders-to-orders-reducedExample output
GROUP orders-to-orders-reduced COORDINATOR-NODE 0 COORDINATOR-PARTITION __consumer_offsets/0 STATE Stable BALANCER cooperative-sticky MEMBERS 1 TOTAL-LAG 0 TOPIC PARTITION CURRENT-OFFSET LOG-START-OFFSET LOG-END-OFFSET LAG MEMBER-ID CLIENT-ID HOST orders 0 3 0 3 0 redpanda-connect-15d7a80f-590f-4cde-bc16-4854fa2754 redpanda-connect 10.0.0.1 orders 1 3 0 3 0 redpanda-connect-15d7a80f-590f-4cde-bc16-4854fa2754 redpanda-connect 10.0.0.1 orders 2 3 0 3 0 redpanda-connect-15d7a80f-590f-4cde-bc16-4854fa2754 redpanda-connect 10.0.0.1 -
Compare the record counts of the two topics. For each topic, the number of available records is the sum of
HIGH-WATERMARKminusLOG-START-OFFSETacross its partitions. A higher count on the new topic means the copy wrote duplicates:rpk topic describe orders -p rpk topic describe orders-reduced -pExample output
PARTITION LEADER EPOCH REPLICAS LOG-START-OFFSET HIGH-WATERMARK 0 0 1 [0] 0 3 1 0 1 [0] 0 3 2 0 1 [0] 0 3 PARTITION LEADER EPOCH REPLICAS LOG-START-OFFSET HIGH-WATERMARK 0 0 1 [0] 0 9Nine records across the three original partitions, and the same nine on the single partition of the new topic.
-
Point your producers and consumers at the new topic. Consumers start from the beginning of the new topic unless you set their offsets with
rpk group seek. -
Stop the pipeline with Ctrl+C.
-
When you no longer need the original topic, delete it to reclaim storage. See Delete a topic.
| Do not delete the original topic until the new topic holds the data you expect and your consumers are running against it. Deleting a topic deletes its data. |
Change the cleanup policy
The cleanup policy determines how to clean up the partition log files when they reach a certain size:
-
deletedeletes data based on age or log size. Topics retain all records until then. -
compactcompacts the data by only keeping the latest values for each KEY. -
compact,deletecombines both methods.
Unlike compacted topics, which keep only the most recent message for a given key, topics configured with a delete cleanup policy provide a running history of all changes for those topics.
All topic properties take effect immediately after being set. Do not modify properties on internal Redpanda topics (such as __consumer_offsets, _schemas, or other system topics) as this can cause cluster instability.
|
For example, to change a topic’s policy to compact, run:
rpk topic alter-config [TOPICS…] —-set cleanup.policy=compact
Configure write caching
Write caching is a relaxed mode of acks=all that provides better performance at the expense of durability. It acknowledges a message as soon as it is received and acknowledged on a majority of brokers, without waiting for it to be written to disk. This provides lower latency while still ensuring that a majority of brokers acknowledge the write.
Write caching applies to user topics. It does not apply to transactions or consumer offsets: data written in the context of a transaction and consumer offset commits is always written to disk and fsynced before being acknowledged to the client.
Only enable write caching on workloads that can tolerate some data loss in the case of multiple, simultaneous broker failures. Leaving write caching disabled safeguards your data against complete data center or availability zone failures.
Configure at topic level
To override the cluster-level setting at the topic level, set the topic-level property write.caching:
rpk topic alter-config my_topic --set write.caching=true
With write.caching enabled at the topic level, Redpanda fsyncs to disk according to flush.ms and flush.bytes, whichever is reached first.
Remove a configuration setting
You can remove a configuration that overrides the default setting, and the setting will use the default value again. For example, suppose you altered the cleanup policy to use compact instead of the default, delete. Now you want to return the policy setting to the default. To remove the configuration setting cleanup.policy=compact, run rpk topic alter-config with the --delete flag:
rpk topic alter-config [TOPICS...] --delete cleanup.policy
List topic configuration settings
To display all the configuration settings for a topic, run:
rpk topic describe <topic-name> -c
The -c flag limits the command output to just the topic configurations. This command is useful for checking the default configuration settings before you make any changes and for verifying changes after you make them.
The following command output displays after running rpk topic describe test-topic, where test-topic was created with default settings:
rpk topic describe test_topic
SUMMARY
=======
NAME test_topic
PARTITIONS 1
REPLICAS 3
CONFIGS
=======
KEY VALUE SOURCE
cleanup.policy delete DYNAMIC_TOPIC_CONFIG
compression.type producer DEFAULT_CONFIG
max.message.bytes 20971520 DEFAULT_CONFIG
message.timestamp.type CreateTime DEFAULT_CONFIG
redpanda.datapolicy function_name: script_name: DEFAULT_CONFIG
redpanda.remote.delete true DEFAULT_CONFIG
redpanda.remote.read false DEFAULT_CONFIG
redpanda.remote.write false DEFAULT_CONFIG
retention.bytes -1 DEFAULT_CONFIG
retention.local.target.bytes -1 DEFAULT_CONFIG
retention.local.target.ms 86400000 DEFAULT_CONFIG
retention.ms 604800000 DEFAULT_CONFIG
Delete a topic
To delete a topic, run:
rpk topic delete <topic-name>
When a topic is deleted, its underlying data is deleted, too.
To delete multiple topics at a time, provide a space-separated list. For example, to delete two topics named topic1 and topic2, run:
rpk topic delete topic1 topic2
You can also use the -r flag to specify one or more regular expressions; then, any topic names that match the pattern you specify are deleted. For example, to delete topics with names that start with “f” and end with “r”, run:
rpk topic delete -r '^f.*' '.*r$'
Note that the first regular expression must start with the ^ symbol, and the last expression must end with the $ symbol. This requirement helps prevent accidental deletions.
Delete records from a topic
Redpanda allows you to delete data from the beginning of a partition up to a specific offset (a monotonically increasing sequence number for records in a partition). Deleting records frees up disk space, which is especially helpful if your producers are pushing more data than anticipated in your retention plan. Delete records when you know that all consumers have read up to that given offset, and the data is no longer needed.
There are different ways to delete records from a topic, including using the rpk topic trim-prefix command, using the DeleteRecords Kafka API with Kafka clients, or using Redpanda Cloud.
|
|
When you delete records from a topic with a timestamp, Redpanda advances the partition start offset to the first record whose timestamp is after the threshold. If record timestamps are not in order with respect to offsets, this may result in unintended deletion of data. Before using a timestamp, verify that timestamps increase in the same order as offsets in the topic to avoid accidental data loss. For example:
|