I’m curious as to which tools and technologies you all are using to keep track of all those services you are deploying, whether it be resource tracking, network traffic, logs, traces, or uptime.

As a bonus question, how have you organized your network or your services to reduce the overhead of implementing observability?

  • faltryka@lemmy.world
    link
    fedilink
    English
    arrow-up
    20
    ·
    3 days ago

    If you’re wanting to go hard in the paint here, Grafana has a great OSS stack with Grafana, Loki, Mimir, Tempo, and Alloy.

    • NotSteve_@lemmy.ca
      link
      fedilink
      English
      arrow-up
      3
      ·
      3 days ago

      Can confirm it’s a great stack! We run it at my work and its quite nice to work with. I’ve been thinking of replacing Prometheus in my homelab with Alloy

          • ℍ𝕂-𝟞𝟝@sopuli.xyz
            link
            fedilink
            English
            arrow-up
            2
            ·
            edit-2
            1 day ago

            Mimir is just the DB, no metric scrapers, but it’s distributed and capable of high volume ingestion, retention and querying. Can’t scrape metrics though.

            Alloy has no storage persistence, so it’s just the scraper part, but is massively scalable, its basically a Prometheus Agent on steroids. Can do logs, traces, Pyroscope profiles too.

            So instead of deploying one server to scrape and store metrics, you deploy one to store and one to scrape.

    • dandiandy@lemmy.caOP
      link
      fedilink
      English
      arrow-up
      1
      ·
      3 days ago

      It looks really compelling. Really flexible with deep options. Seems like I could spend 10 years building the dashboard alone!

      • foggy@lemmy.world
        link
        fedilink
        English
        arrow-up
        1
        ·
        3 days ago

        I stood up grafana at work, and I make dashboards to keep my role secure.

        Everyone uses my shitty little dashboards :)