Writing / 2023

Claude vs GPT-4: Two Weeks Using Both

Claude and GPT-4 both went public on March 14. How they compare on instructions, refusals, and reasoning, and why teams shouldn't marry one model.

I’ve been using Claude and GPT-4 side by side since both went public on March 14. Two weeks isn’t long, but it’s enough to see different personalities and different tradeoffs. This is how they compare for someone who cares about the output, not the branding.

The Constitutional AI Thing

Anthropic trains Claude using what they call Constitutional AI. Instead of relying purely on human raters, they define a set of principles (be helpful, honest, and harmless) and train the model to critique and revise its own outputs against those principles.

The interesting part for me is that the method is published. The Constitutional AI paper lays out the approach and the principles used in the research. OpenAI has its own alignment approach, but it’s less transparent about the specifics. Whether that matters to you depends on how much you care about being able to reason about why a model behaves the way it does.

Where Claude Wins

Claude is noticeably better at following nuanced instructions. When I ask for a specific format with specific constraints, Claude tends to honor the full request more consistently. GPT-4 sometimes interprets instructions loosely, especially length constraints.

Claude also refuses less aggressively on borderline requests. GPT-4 has a hair trigger on anything that looks remotely sensitive, which is annoying when you’re trying to do legitimate work in domains like security or healthcare. Claude is more calibrated: it will engage with difficult topics while still maintaining guardrails.

For longer conversations, Claude holds context better in my experience. Less drift, fewer moments where it seems to forget what we discussed three messages ago.

Where GPT-4 Wins

Raw reasoning power. When I throw a genuinely hard problem at both (complex code logic, multi-step analysis, ambiguous tradeoff evaluation), GPT-4 produces stronger outputs more reliably. The gap is real.

GPT-4 also has a larger ecosystem. More integrations, more tooling, more community knowledge about prompting patterns. That matters when you’re building production systems and need to solve problems fast.

And GPT-4 accepts images, at least on paper. OpenAI showed image input at launch but hasn’t opened it to the public yet. Once it ships, image understanding opens up product surfaces that text-only models can’t touch, and Claude has nothing comparable announced.

What This Means for Teams

If you’re building AI features, don’t marry a single model. Keep a thin abstraction layer so you can swap providers without a rewrite. Test your prompts against both and pick the one that performs better for each specific use case.

Some practical habits:

  • Test prompts that probe safety boundaries . Log the refusals. Know where each model draws the line for your use case.
  • Write down your behavioral expectations in the same repo as your app code. This is your spec, regardless of which model sits behind it.
  • Accept that the tradeoffs are real. A cautious model will refuse some legitimate requests. A flexible model will occasionally let something through that it shouldn’t.

Competition between Claude and GPT is good for everyone building on top of these models. Different approaches, different strengths, and pressure on both teams to improve.

References