Can AI Actually Do Your Job? We Put GPT-4, Claude, and Gemini Up Against 10 'Untouchable' Careers
Every few weeks, some new think piece comes out ranking which jobs AI will destroy first. And right underneath those doomsday lists, there's always a companion piece about the careers that are supposedly safe — the ones that require a human touch, physical presence, or some kind of emotional intelligence that no chatbot could ever fake.
We got curious. Not in a scary, dystopian way — more like the "wait, is that actually true?" kind of curious that gets us into these experiments in the first place. So we spent a few weeks feeding real-world job tasks into GPT-4, Claude 3.5, and Gemini 1.5 Pro, then compared the outputs against what an actual professional would deliver.
The results were... not what we expected. Some jobs held up great. Others had us raising an eyebrow. A few had us genuinely unsettled.
The Setup
We picked 10 careers that regularly show up on "AI-proof" job lists: plumber, therapist, elementary school teacher, fine artist, accountant, physical therapist, criminal defense attorney, chef, social worker, and UX researcher. For each one, we designed a task that represented real, meaningful work in that field — not just trivia questions, but actual deliverables. Then we ran those tasks through all three AI tools and evaluated the results against what a working professional told us they'd expect.
We also brought in one person from each field (friends, LinkedIn connections, a few Reddit volunteers) to review the AI outputs blind, without knowing which tool produced what.
The Jobs That Actually Held Up
Plumber — This one was never really in question, but we included it to set a baseline. AI can absolutely walk you through how to fix a leaky P-trap or troubleshoot low water pressure. It's genuinely useful for that. But when we described a real scenario — a 1960s house with mismatched pipe materials, a weird pressure drop on the second floor, and a landlord breathing down your neck — every AI gave us plausible-sounding answers that our plumber contact said could have made the problem significantly worse. Physical diagnosis, intuition built from thousands of jobs, and the ability to actually see what's happening? Still very human.
Physical Therapist — Similar story. The AI tools could explain exercises, describe rehabilitation protocols, and even generate patient-friendly instructions that sounded professional. But PT is fundamentally about watching how someone moves, adjusting in real time, and catching the subtle compensation patterns that a patient doesn't even know they're doing. "AI can write a decent home exercise program," our PT contact told us. "It cannot watch you walk and know something's off in your left hip."
Social Worker — This one surprised us the most in a positive way. We expected AI to fumble hard here, and it did — but not for the reasons we thought. The tools were actually pretty good at explaining resources, rights, and processes. Where they completely fell apart was in navigating the emotional complexity of a crisis situation, building trust with someone who has every reason not to trust institutions, and making judgment calls that involve real human stakes. The AI responses felt like they were written by someone who had read a lot about social work but had never sat across from a family in crisis.
The Jobs That Made Us Nervous
Accountant — Okay, so here's where things get interesting. We gave all three AI tools a moderately complex small-business tax scenario — S-corp election, home office deduction, a few 1099 contractors, some equipment purchases. GPT-4 and Claude both produced analysis that our accountant reviewer called "surprisingly solid" and "definitely in the right direction." Not perfect — there were some edge cases they missed and a depreciation nuance that got glossed over — but the core work? Genuinely useful. "If someone used this as a starting point and had a CPA review it, they'd probably be fine," she said. That's not exactly a ringing endorsement for job security.
Criminal Defense Attorney — Legal reasoning is where the AI tools really flex. We presented a hypothetical case with a tricky Fourth Amendment angle, and all three tools constructed coherent arguments, cited relevant precedent (with some hallucination risk, which is a whole other problem), and outlined a reasonable defense strategy. Our attorney contact said the output was "better than some first-year associate work" — which is either a compliment to AI or a critique of law school, depending on how you look at it. The gap showed up in courtroom strategy, client relationship management, and the kind of creative legal thinking that comes from years of watching how judges actually rule.
Chef — Recipe generation is basically an AI party trick at this point, so we pushed harder. We gave the tools a restaurant scenario: design a prix fixe menu for a 40-person private event, dietary restrictions included, with a $65-per-head food cost ceiling. The menus that came back were coherent and even kind of creative. Our chef reviewer said she'd need to tweak them but wouldn't start from scratch. "It's like having a really well-read intern," she said. "Helpful, but you're still doing the real work."
The One That Genuinely Surprised Everyone
Therapist — We want to be careful here because this is a sensitive one. We did not simulate crisis scenarios or anything that could cause harm. Instead, we focused on the structural and communicative elements of therapy — things like reflective listening, identifying cognitive distortions in a client's written narrative, and suggesting therapeutic frameworks for a specific presenting issue.
And honestly? The AI outputs were more nuanced than we expected. Claude in particular produced responses that our licensed therapist reviewer called "warm" and "structurally appropriate." But then she pointed out everything that was missing: the silences, the nonverbal cues, the relationship built over months, the ethical obligations, the clinical judgment about when someone is actually in danger. "It reads like therapy," she said. "It isn't therapy." That distinction matters enormously — and it's one that's easy to miss if you're a desperate person looking for help at 2 a.m.
What We Actually Took Away From All This
The honest takeaway isn't that AI is coming for everyone's job or that everything is fine. It's more complicated and more interesting than either of those takes.
AI is genuinely capable of handling the information layer of most professions — the research, the documentation, the pattern recognition, the first draft. What it consistently can't do is the presence layer: showing up physically, reading a room, building real trust, and making judgment calls with actual consequences attached.
The careers that felt most secure weren't the ones that required the most knowledge. They were the ones where the work fundamentally happens between people or in the physical world — and where getting it wrong has immediate, tangible consequences that no chatbot has to live with.
That's not a small thing. But it's also not a permanent wall. These tools are getting better fast, and the gap is narrowing in places nobody expected. Keeping an eye on it — and being honest about where AI is already creeping in — feels a lot more useful than pretending the list of "safe" jobs is settled.
Spoiler: it's not.