Latest

  • The estimated reading time for this article is about 3 minutes.

    PB3K architectural review pipeline

    Whelp, that certainly is a flowchart. Let's stick a pin in that for a bit.

    I have a modest unified memory archive AMD Strix Halo machine with 64GB of RAM running Ubuntu with ROCm. I bought this in early September 2026 and with Amazon points, it was still too much money. Certainly, the most I have ever spent on non-Apply hardware.

    I use opencode (and have a basic Go subscription) and ChatGPT (with a Plus subscription). My goal is to minimize my use of frontier LLMs for private projects. I think now have such a rig.

    I run three models on this host (called Colossus):

    • Qwen3.8-27B-UD-Q4_K_XL (a dense model with a 64K context)
    • Qwen3-Coder-30B-A3B-Instruct-Q4_K_M (a Mixture of Experts model with 128K context)
    • gpt-oss-20b-F16 (MoE model with 128K context)

    Using the frontier model, benchmarking scripts were created to dial in the settings on these LLMs.

    Only one LLM runs at a time, but I have a script that launches llama server with a router. I run that one script and then I can select the model I want to use in OpenCode.

    The MoE models (once they are loaded) run very well and have snappy token generation. However, their reasoning is questionable. Their tool use is OK.

    The Qwen 3.8 model is slow (about 10 t/s) but accurate. This is my go-to model when I need the LLM to do software engineering and discovery tasks correctly.

    Which brings me to this monster flow chart.

    I have a game called ProspectBoy 3000. And although a lot of the game works as expected, the architecture went a little pair-shaped in places. I have fixed some of this by hand, but know I constructed an architectural review with Qwen.

    This workflow is particularly challenging because there is only one agent in the loop at a time. Therefore, a perl driver is used to launch opencode with instructions that examine only one Perl library file at time. The results are written to disk. The list of modules to examine is also updated.

    This review has been going for 16 hours without my intervention. I will likely go for another 48 hours.

    Qwen is fantastic. I can work around latency. I can't work around inaccuracy.

    When I need interactivity, I have the MoE models which are great too.

    Given the insane prices of VRAM, I have to believe 2027 will be the year of local LLMs. The good news is, many workflows already are possible on 64GB machines. 128GB is a luxury as is speed. If you can get it though, take it.

    Cheers,

  • til chatgpt can read atari2600 roms

    ChatGPT can read atari2600 ROMs. I uploaded Adventure, told it that I think there is bug
    in 3rd variant of the game where items get put into inaccessible places, and the LLM
    confirmed the bug!

    This variant distributes items by using the table the statically placed items in
    various location in variant 2. It takes that table and simply shuffles the locations
    assigned to those items. So one set of locations have a fixed place where item will
    occur. It's just that items get distributed differently to those location when you
    reset a variant three game. Super-slick hack!

    There is no validation logic, however, that prevents the gold key from starting
    in the gold castle. Wump, wump. Start again.

    Anyway, I don't think Warren Robinett is accepting patches.
  • The estimated reading time for this article is about 4 minutes.

    Deep Learning (DL) is a slightly out of date term for training neural networks on digital content in order to recognize patterns. This technique led to driver-assisted cars, medical insights, self-focusing cameras, and a host of other technologies and produce affordances we have all come to depend upon.

    I do not hear a lot of complaints on the Internet that self-focusing cameras are running photography.

    The marketing term "generative Artificial Intelligence" stuck in my increasing graying maw from day one. Large Language Models (LLM) are still primarily DL engines. Sure, there is some "reasoning" sprinkled on top. That word "reasoning" should be looked at like a legal term, rather than understand as a plain English term.

    The best way to look at the output of LLMs is a bit like the old Plinko game. You put a disk (or a your word salad) on the top and that bounces around randomly until you get something out at the bottom of the board. The AI tech companies are trying to make the shape of that Plinko board "useful," to which I say "good luck."

    LLMs are another in a long lime of tech tools that over-promise capability and under-promise value. The only ones who think "the game has changed" are those same managers who have been looking to fire employees for my entire career. AI has only made their fever dreams more sweaty.

    Lest you think I am a Luddite advocating for the pitchforks and bonfires, please understand that I use AI tools every day gladly. However, just like with a chainsaw, I do not let the AI tools think for me. I have to drive, cajole, plead, and sometimes override the code tools like Claude Code and OpenCode produce. That is my value as an engineer and it has been for decades. I am responsible for the architecture of an application. Sometimes (often) that architecture is not just the sweeping design document, but the crafting to the way smaller parts handle just enough work to be useful, but not so much that they become monolithic, over-balanced clown shows.

    AI does not do this last part well. It cannot.

    By now, I trust we can all recognize all but the best LLM generated voice messages. We can easily spot AI crafted text. This combo is the bane of YouTube right now. However, I will admit that AI-generated graphic appeal to me as a non-graphic artist. If the image has no factual mistakes, I am hard-pressed to tell the difference. That is certainly not art, but neither are stock photos and clip art.

    The danger of AI is that people can become intellectually lazy. Rather than have an LLM suggest copy edits for text, user copy the AI output blindly. But we have been here before. Twenty five years ago, shame and opprobrium where heaped on those who "googled" for answers. Thirty years ago, movies and TV that used computer generated graphics were similar denigrated.

    Creative and curious people will continue to do great things using the tools that are available to them at the time, including LLMs. Lazy people will continue to be lazy, just at a faster pace. As I said, this has all happened before and should not overly concern you.

    As for AI-slop, that is a curious category of output. I suspect as we demand higher-quality content, AI-slop will become unprofitable. But low-quality content has always been with us. Media with factual errors is the foundation of my childhood.

    To quote Battlestar Galactica one more time: this has all happened before.

    This essay is in response to my friend Nate's very understandable complaints about lazy people using AI to be unhelpfully lazier.

  • The ProspectBoy 3000

    I have a love for the awkward teenage years of the Internet (circa 1995 - 2005). I wanted to recreate an echo of that time for those who missed it. That was the period before "content" and "monetization" were words that ordinary users cared about. People made stuff of the web that they hoped people might like or find useful or did it just because they could.

    In that spirit, I have built a BBS door-like game called ProspectBoy 3000:

    https://games.taskboy.com/pb3k/

    It takes place centuries after our world has been forgotten. You are one of many prospectors who have gathered around a strange object called the Magic Mountain. It's not a natural mountain, but an enormous structure that contains mysterious artifacts that can be made useful by your cunning skills. Sell your wares to agents of the five factions vying to rule. Your actions help shape the world. Hurry! The mountain disappears every 30 days only to reappear miles away for the whole process to start again.

    The game is free to play and open source. This isn't a revenue play. It isn't a resume builder. It's a labor of love. I hope you enjoy it.