• kungen@feddit.nu
    link
    fedilink
    arrow-up
    7
    ·
    24 hours ago

    My team made the “mistake” that all new projects need to be built and approved on our dev clusters before we allow them to be deployed to the prod clusters. We’re very explicit that dev is ephemeral, and that everything should be as code, so everything heals itself…

    Despite this, whenever we reprovision the dev env, we always get emails from several different people saying that we wiped their prod databases and such. It’s crazy.

      • kungen@feddit.nu
        link
        fedilink
        arrow-up
        3
        ·
        21 hours ago

        We don’t usually do it for fun, but I’m sure my teammate who has petitioned deploying a “chaosmonkey”-thingy to our prod clusters would love it 😁

        • pivot_root@lemmy.world
          cake
          link
          fedilink
          arrow-up
          4
          ·
          edit-2
          21 hours ago

          Your teammate has the right approach here. People will quickly learn to stop relying on the dev environment in production deployments when the environment becomes unreliable.

        • flambonkscious@sh.itjust.works
          link
          fedilink
          English
          arrow-up
          2
          ·
          20 hours ago

          I really wish we could get to that level of maturity (cries in healthcare)

          Not sure if we should take the approach of biting the bullet and forcing it on ourselves or slowly growing to it (under a hostile govt that sacked ~30% of the countries ICT team)

          • Peffse@lemmy.world
            link
            fedilink
            arrow-up
            1
            ·
            16 hours ago

            back when my team had proper coverage, we used to do chaosmonkey style tests for hosted health systems. We’d schedule an event with the hospital IT staff. Say… 1 hour on a specific day near midnight. The local IT staff would put out a notice that for the hour the system may be inaccessible. Then we’d start a test suite, as ungracefully as possible. We’d run a abort shutdown command on one server, see how RHCS would handle it. Bring it back up stable. Silo a database instance. See how that handled. Yank out the Red1 cable from the switch, see how the network balanced the new traffic over Red2/Blue. That kind of stuff. We caught a few problems that way.

            The clients hated it, but it really made the system more bulletproof for when a new recruit ran a command in the wrong window and took down an unrelated resource.