The M4 Got Me

I never buy first-gen Apple products. The first Intel Macs were bad. You wait for the .1, the .2, sometimes the .4 before things get solid. That was my rule until the M4 MacBook Pro Max broke it. People were losing their minds about running AI models locally and I wanted in.

I have a love-hate thing with Apple. The hardware runs forever. The interfaces are streamlined in ways nobody else matches. But you need a bag of dongles to plug in a USB stick or an HDMI cable, and they’ve been iPadifying the MacBook Pro for years. Still, if any company could make on-device AI work, it was the one that controls the entire stack top to bottom.

Apple Intelligence shipped and it was Siri with a new label.

270 Million to 30 Billion Parameters

Ollama made pulling models dead simple. LM Studio too. I tested everything from 270 million parameter models to 30 billion parameter ones on the new machine. Got to know their quirks, their strengths, where they break. A 7B model has a specific personality that’s nothing like a 13B, and a 30B model gives you a taste of intelligence that makes its limitations sting more. That exercise was worth the price of the hardware by itself.

The ceiling was obvious though. We got used to Claude and GPT. Anything running on a laptop felt like going back to early chatbots. Not quite that bad, but the gap was wide enough to notice on every prompt.

I tried Pinokio for running containerized model apps. Cool concept, flaky in practice.

I have ADHD. Keeping tasks organized across work, personal projects, life stuff is a mess. My most reliable system is a paper notepad. I figured a local model hooked into Apple’s mail and reminders could triage things for me. A Siri that understood context. The moment any cognitive task showed up, the local model fell over. Understanding intent, prioritizing, making judgment calls. Every interesting part needed an API call to a bigger model somewhere else. Defeats the purpose of running things locally.

Models Getting Smarter

Gemma 2 was forgettable. Gemma 3 changed things. Structured outputs got noticeably better. Google’s AI Studio had a barista demo I spent real time trying to jailbreak. I did break it, but it took more effort than I expected. Impressive for a model that runs on a laptop.

A small model fine-tuned for a narrow task can be excellent at that task. It won’t debug your codebase, but it’ll stay in character as a barista better than models ten times its size. That was the shift in my thinking. General intelligence isn’t the only game.

Gemma 3N pushed this further. On-device, multimodal, with effective parameter counts well below the actual model size. The E2B variant operates with the memory footprint of a 2 billion parameter model despite packing 5 billion parameters total. And it could do function calling. A model that small, calling functions consistently, was new. There’s been a whole leaderboard tracking which models can call functions well, but most of those were big models. Gemma 3N showed it was possible at the edge.

Apple Foundation Models arrived after that. 3 billion parameters, quantized to 2 bits, built for Apple’s platform. Not trying to be Claude. Apple built something that works within their walled garden, tool calling included, and shipped it.

Soft-Coded Logic

This is where I think things are heading. Models like Gemma 3N and Apple’s Foundation Models can do tool calling. Not perfectly. Not every time. But the direction is clear.

The idea: instead of hard-coding all your logic, you let the model handle some of the routing. User says something, model figures out which function to call, function does the work. You’re not replacing your entire codebase with an LLM. You’re using it as a piece of the logic, handling the parts where intent is fuzzy and a wall of if-statements would be brittle.

You don’t need Claude-level intelligence to check email or add a reminder. Those tasks need routing, not reasoning. A 3 billion parameter model that’s decent at tool calling opens up a category of apps that used to require either hardcoded conditionals or API calls to smarter models. Add a cron job and you have something like a local, offline personal assistant that doesn’t phone home. That ADHD task manager I failed to build a year ago? It’s getting feasible. Not because the model got brilliant, but because tool calling lets a limited model do useful things.

Apple says it themselves: the on-device model “is not designed for world knowledge or advanced reasoning.” Fine. If it can call the right function with the right parameters most of the time, that’s a different kind of useful than what we’ve been chasing.

The Side Quests

Beyond LLMs, I’ve been running music generation models, image generation with Flux, and voice cloning with Qwen3 TTS. That last one cloned my voice from a few seconds of audio. My speaking speed, my cadence. Cool and fucking terrifying.

Every YouTube video about running models locally is like an old rom-com. They show you the wedding, never the marriage. A model generating text in a terminal window, look how great. Then you try to build something with it and 80% of the tasks you care about fall flat.

February 2026, it’s still early. But two things are happening at once: models fine-tuned for narrow tasks are getting genuinely good, and tool calling is changing how software gets built. Most of the YouTube demos miss both.

ai, tools, opinion