I’m curious as to which tools and technologies you all are using to keep track of all those services you are deploying, whether it be resource tracking, network traffic, logs, traces, or uptime.
As a bonus question, how have you organized your network or your services to reduce the overhead of implementing observability?


If you’re wanting to go hard in the paint here, Grafana has a great OSS stack with Grafana, Loki, Mimir, Tempo, and Alloy.
Can confirm it’s a great stack! We run it at my work and its quite nice to work with. I’ve been thinking of replacing Prometheus in my homelab with Alloy
It’s curios if you used Prometheus as well
Mimir and Alloy are two specialised forks of Prometheus
How are they specialized?
Mimir is just the DB, no metric scrapers, but it’s distributed and capable of high volume ingestion, retention and querying. Can’t scrape metrics though.
Alloy has no storage persistence, so it’s just the scraper part, but is massively scalable, its basically a Prometheus Agent on steroids. Can do logs, traces, Pyroscope profiles too.
So instead of deploying one server to scrape and store metrics, you deploy one to store and one to scrape.
It looks really compelling. Really flexible with deep options. Seems like I could spend 10 years building the dashboard alone!
I stood up grafana at work, and I make dashboards to keep my role secure.
Everyone uses my shitty little dashboards :)