We resolved the performance bottleneck. All APIs are responding as expected as of now.
The issue has been resolved.
The incident has been resolved.
All latencies are back to normal
We scaled our ingestion and are processing data in time.
Upgrade successfully resolved.
We scaled our infrastructure and are processing media uploads without errors.
This issue has been resolved.
The underlying incident with our infrastructure (https://statuspage.incident.io/clickhousecloud/incidents/01KT1G25S9PBKM7VJEB146680G) has been resolved.
We've fully caught up and process LLM as a judge in realtime again.
We have scaled our infra to drain the queue backlog. Evals are executed without delay again.
The queue backlog has been fully processed for legacy trace and dataset targeted evals. There is no longer a delay for eval executions.
We observe full recovery across all API and UI routes.
Latencies and error rates have recovered.
Maintenance has completed
The underlying issue was resolved.
We resolved the underlying issue. A distributed cache roll out for ClickHouse caused the latency and error spikes. The roll out was reverted and we see normal API performance as of now.
We caught up with our queues and have regular processing times again.
All data is processed in time.
We've reverted a bad patch and see recovery. All routes should work as expected.
·