Tag: big-data
All the articles with the tag "big-data".
Kafka (Part 5): Consumer Callbacks, Scheduled Retries, and Rebalancing
This largely completes a messaging system built with Kafka. The key code has been explained across the articles; see the source repository for the complete design and implementation. The next Kafka article is planned to cover synchronizing Oracle data to PostgreSQL with Kafka Connect.
Kafka (Part 3): Sending JSON with a Shared Serializer and Improving Producer Throughput
Implements a Kafka producer that sends JSON messages, using a shared serializer for different objects and configuring asynchronous sending, failure recording, and retries. Based on actual email-message sizes, it discusses how compression, batch size, request size, and waiting time affect throughput.
Kafka (Part 1): A Single-Node KRaft Setup with Docker Compose, Kafka UI, and Prometheus JMX Exporter
Builds a single-node Kafka environment in KRaft mode with Docker Compose and integrates Kafka UI and Prometheus JMX Exporter. It explains listeners, roles, memory, and monitoring configuration, then covers startup verification, JMX port conflicts, and resource-related troubleshooting.
Using Solr (Part 2): Full and Incremental Imports from MySQL
Explains how to configure Solr DataImportHandler for full and incremental imports from MySQL and periodically refresh data with operating-system scheduled jobs. It also covers tinyint fields, multivalued fields, and Transformers, and outlines DIH components and changes in later versions.