The Bottleneck: 45-Minute Batch Lag in Financial Operations
A fintech analytics client collected telemetry, transaction requests, and risk signals from across dozens of customer applications. As daily volume crossed 30 million events, their traditional relational warehouse buckled under the load.
Nightly batch ETL cron jobs frequently failed or ran into the morning business hours. Operations teams were making critical risk decisions on data that was up to 45 minutes stale, while complex multi-table joins took over 30 seconds to execute.
“In financial risk intelligence, 45-minute-old data is functionally useless. We needed sub-second freshness and millisecond query speeds without exploding compute costs.”
Production Performance Benchmarks
Daily Ingested Events
Sustained stream processing with auto-scaling ingestion workers.
p99 Query Latency
Complex analytical rollups computed directly against columnar storage.
Dropped Payloads
Guaranteed exactly-once semantics even during 15k req/s traffic spikes.
Storage Optimization
ZSTD columnar compression reduced storage footprint by over 80%.
The Stream Architecture: Decoupled Ingestion & Columnar Storage
Plexel designed a zero-loss streaming pipeline combining lightweight Rust ingestion edge daemons, partitioned Kafka message brokers, and an optimized ClickHouse analytical engine:
1. High-Concurrency Rust Edge Ingestors
Lightweight ingestion microservices written in Rust with tokio async runtimes to terminate inbound client traffic. They validate JSON schemas, strip malformed payloads, and batch messages into Kafka partitions with microsecond overhead.
2. Partitioned Kafka Event Backbone
Events partitioned by tenant and customer ID to guarantee strictly ordered delivery per stream while allowing parallel downstream consumption across independent worker pools.
3. Direct-to-Columnar Streaming (ClickHouse)
ClickHouse replaced the traditional warehouse. Using ReplacingMergeTree and AggregatingMergeTree engines with materialized views, incoming events are indexed and aggregated on the fly as they land on disk.
4. Periodic Spark Reconciliation
Apache Spark was retained solely for asynchronous, non-blocking daily auditing and deep historical anomaly detection without degrading real-time query engines.
From Stale Batches to Real-Time Decisioning
Within 8 weeks of deployment, data freshness plummeted from 45 minutes to under 600 milliseconds. Leadership and operational teams interact with live telemetry dashboards with instantaneous response times across billions of historical records.
p99 query latency on multi-million row aggregations.
Pipeline uptime through peak transaction surges.
Storage reduction via modern columnar compression.