• Gil, The Shitposter Of Ages@lemmy.dbzer0.com
    link
    fedilink
    English
    arrow-up
    15
    arrow-down
    11
    ·
    15 hours ago

    I could either take several tedious hours to write out several thousand regex patterns by hand to automatically recognize differing formats of notation for this thing I’m doing IRL…

    or I could have an LLM do it in a few seconds, and it’ll get 99.99% of the cases, missing maybe one case every hundred thousand…

    take your guess as to which option I chose.

    LLMs are a tool. I use it as such.

      • OfCourseNot@fedia.io
        link
        fedilink
        arrow-up
        6
        ·
        13 hours ago

        write out several thousand regex patterns

        Oh my! This thing is more likely to gain sentience than any llm.

        • Gil, The Shitposter Of Ages@lemmy.dbzer0.com
          link
          fedilink
          English
          arrow-up
          3
          ·
          12 hours ago

          it is a truly absurd number, simply due to the inane BS variations on terminology and notation.

          It might have been more reasonable to train a local LLM on this material and then just query -that-… but hey 30ish seconds of claude’s free tier usage; and i’ve got a script that’s ~99.99% as good as a human in terms of pattern recognition. Pretty neat when you think about it.

          • OfCourseNot@fedia.io
            link
            fedilink
            arrow-up
            1
            ·
            6 hours ago

            And ‘truly absurd number’ might still be a bit of an understatement here. Saying that Claude can generate several thousand regexes for you is not really a selling point for llms in my book.

            On the other hand. Have you tried running them on the bible or something like that? I bet it finds it talks about ww2, 9-11, the nintendo wii, and shit…

            • Gil, The Shitposter Of Ages@lemmy.dbzer0.com
              link
              fedilink
              English
              arrow-up
              1
              ·
              edit-2
              4 hours ago

              If, by bible, you refer to the holy book of the local Building Code; then no, I have not.

              and again; the regex in the script generated works, it’s just a teeny bit cursed and maybe a little bit sentient. Praise be to the Omnissaiah, and its blessed machine.

      • mobyduck648@lemmy.world
        link
        fedilink
        arrow-up
        6
        ·
        15 hours ago

        Can’t be worse than when I was an intern, the product was literally based around parsing arbitrary HTML with regex and not even decent regex, POSIX regex from like the '80s for reasons known only unto the programming gods.

      • ZILtoid1991@lemmy.worldOP
        link
        fedilink
        arrow-up
        3
        ·
        15 hours ago

        When I wrote my numerical parser, people told me to use regexes to detect what kind of numerical value is being parsed. I just did it in one go. Only regex-like checking is for ISO datetime values.

        • Gil, The Shitposter Of Ages@lemmy.dbzer0.com
          link
          fedilink
          English
          arrow-up
          1
          ·
          12 hours ago

          regex is computationally pretty cheap, from what I’ve read; due to something about it being a simple AND gate or something… . IDR, and to be honest it doesn’t really matter in the era of 16-core/32thread 5GHz processors… but when you just want it to work… it does “just work”. /shrug.

      • Gil, The Shitposter Of Ages@lemmy.dbzer0.com
        link
        fedilink
        English
        arrow-up
        3
        arrow-down
        1
        ·
        14 hours ago

        Not really. There will only ever be about a hundred people who ever use this tool, and it’s really just a massive timesaver that keeps people from needing to open a smorgasbord of files to get a few tidbits from each file; then collate it into one doc…

        If I get ~99.99% of all the information, and one hundred people use this, odds are, nobody will ever see the .01% edge case of missed info. and even if it does happen, it’s not like it’s obsfucated information in each file, it’s just a pain to open a hundred files and wait for them all to load at once through the disk-heavy GUI application we have to use.

        simpler to just read via the regex patterns, and then put it out into the terminal for the info needed.

    • jj4211@lemmy.world
      link
      fedilink
      arrow-up
      3
      arrow-down
      1
      ·
      10 hours ago

      Frankly, your use case sounds like something has gone horribly wrong from the outset. If something needs several thousand regex patterns, then something is very very wrong, and there’s zero chance you have a comprehensive solution as it stands. Either trying to use regex for a use case that some natively AI approach might actually be warranted for, or regex is being used poorly and each one is too limited, or some other approach entirely is warranted.

    • NotSteve_@lemmy.ca
      link
      fedilink
      arrow-up
      7
      arrow-down
      1
      ·
      14 hours ago

      or I could have an LLM do it in a few seconds, and it’ll get 99.99% of the cases, missing maybe one case every hundred thousand…

      Even if the LLM does things perfectly 100% of the time, how do you expect to maintain your skills if you offload all of your work onto a chatbot? What happens when you start offloading more and more until you’re not actually doing anything except typing in prompts?

      • Gil, The Shitposter Of Ages@lemmy.dbzer0.com
        link
        fedilink
        English
        arrow-up
        1
        arrow-down
        1
        ·
        13 hours ago

        My skillset is geared towards meatsacks and things that go boom. LLMs have been a massive boon to the 1% of my job that needs a code monkey.

        As long as people still talk to each other, and things need exploding, I’ll have a job. Heck, if we somehow get to a point where money isnt needed, I’d still do what I do simply 'cuz I like it.

    • mobyduck648@lemmy.world
      link
      fedilink
      arrow-up
      4
      arrow-down
      2
      ·
      14 hours ago

      Yeah I’m not a fan of slop merchants and PRs the size of battleships but anyone solving merge conflicts by hand because they hate LLMs is being a tad masochistic in my opinion. I’m in two minds about how much coding itself ought to be automated, but I’m a strong advocate of automating the tedious donkey work we’d absolutely have got rid of years ago some other way if possible.

      It’s also good for avoiding tarpits, sometimes I literally don’t want to learn something because it’s too expensive relative to the value of retaining it. I would rather crawl over a mile of broken glass and hot coals than learn the detailed ins and outs of WSDL just to support a legacy feature which is firmly in the crosshairs of my deprecation gun, to use a totally neutral and not driven by personal bitterness example.

    • makeshift0546@lemmy.today
      link
      fedilink
      arrow-up
      2
      arrow-down
      1
      ·
      14 hours ago

      Lots of folks here create absurdly high expectations that most coders don’t meet. They are out there but they ain’t average. Getting it 80% correct with good human reviews to catch the trash and using it to figure out what’s breaking quickly vs tracking down logs has been great for our team. I laugh at some of the garbage it spits out and the dumb way you can see shit like, ‘I have to be honest I did zero testing’.