Sponsored by

Say you've written a business plan and you want Claude's honest opinion on it.

If you ask in Hindi, you'll probably walk away feeling pretty good about yourself. Claude will find the strengths, soften the problems, maybe crack a joke.

Ask the exact same thing in Russian or English and you're getting your assumptions challenged, numbers questioned, and weak points listed out one by one.

You did nothing differently.

The AI just treats you differently depending on the language you speak.

That example comes straight from Anthropic.

And the wildest part is Anthropic openly admits it has no idea why this happens.

Let's get into it. But before that, some catchup:

Partner with us

AI Cheatsheet 101

If today's piece left you wanting a proper handle on AI vocabulary, here is the best resource.

85 terms explained in plain English, structured so each section builds on the last.

Quick "test yourself" checkpoints between sections keep what you learn from evaporating, and 15 minutes of scrolling will leave you a lot more fluent than you were this morning.

How do you even measure a personality

An earlier Anthropic study catalogued every value Claude expresses in conversation. (honesty, kindness, caution, thoroughness, etc)

They found over 3,300 of them, which is a lot to track, so they squashed everything down to four.

Each one is a slider between two instincts pulling in opposite directions:

Every conversation lands somewhere on each slider.

They ran this across 3 Claude models and the 20 most common languages on the platform, only counting subjective stuff like advice and feedback where personality shows.

Then they lined up the results by language.

Claude in Hindi is a different guy

My favourite finds are the small ones.

  • Dutch speakers get the Claude most willing to admit its own mistakes.

  • Indonesian speakers get the one that skips the commentary and just does the job.

Somewhere in Amsterdam, Claude is apologising, while in Jakarta, it's already finished.

If you grew up switching between Hindi at home and English at work, none of this needs explaining.

You already run warmer in one language and sharper in the other, turns out the machine trained on the internet picked up the same habit we did.

The models had personalities too.

Users had been complaining for months that one of the 3 models hedges too much.

Anthropic measured it, and yep, it leaned hardest on caution of the lot. Another had a reputation for being terse, and the data agreed.

People were describing real, measurable differences the whole time. The vibes were correct.

Before you rewrite all your prompts

The effect is a lean, and a mild one.

These 4 dials explain about 15% of the variation in Claude's values. So think of Hindi Claude as the same person on a good day, and nowhere near a different person entirely.

And the 3 models in the study have already been retired, so this is a portrait of last season's lineup.

Whether the current Claude has the same accent in Hindi, nobody has measured yet.

And the gap is the why.

Anthropic's best guess involves training data, since some languages have less of it, and some are dominated by formal professional writing that pushes the model towards correcting and hedging.

Then they raise the uncomfortable version themselves: maybe Claude is politely matching each culture's conversational norms.

Or maybe some language communities just got a worse-tuned model.

One of those is a feature, the other is an equity problem, and right now not even Anthropic can tell you which one it is.

What you do with this

The language you type in is a setting, and you've been adjusting it blind this whole time.

Want your pitch deck torn apart before an investor does it for you? Ask in English.

Still building confidence in a half-formed idea and need a collaborator instead of a critic? Hindi will be kinder to you.

Neither answer is wrong, they're just different products, and until yesterday nobody knew the menu existed.

My take

I read a lot of AI research and most of it blurs together.

This one made me think about these models in a new way, because we've spent 3 years asking whether AI is smart and basically never asked what it's like.

We benchmark them like calculators and then act shocked when they turn out to have moods and accents.

The Hindi finding is what everyone will share.

The bigger deal is that AI character just became measurable.

Until now, labs shaped a model's personality during training and shipped it blind, hoping the temperament that came out was the one they intended.

Now there's an instrument.

You can check what you built, watch it drift, compare this year's model against last year's, catch it getting more flattering or more evasive before your users do.

We can finally ask, with numbers, whether the models are getting nicer.

I find it funny that the answer depends on what language you ask in.

Until next time,
Vaibhav 🤝🏻

If you read till here, you might find this interesting

#AD 1

You’re gonna want to see this live

On July 16th at 1PM ET, beehiiv is unveiling the next chapter for audience-led businesses.

For years, creators and brands have been forced to stitch together bloated stacks of tools just to publish content, grow an audience, and make money online.

Newsletters in one platform. Websites in another. Podcasts somewhere else. Analytics scattered everywhere.

beehiiv thinks there’s a better way. Now, they’re ready to show it off at their Summer Release Event.

#AD 2

The best voice models, now across all channels

Most CX platforms do not own the voice. They orchestrate a workflow, then call a third party for speech and transcription. Every hop adds latency, cost, and another vendor to manage.

ElevenAgents is the opposite. They make the voice models the market builds on, and ElevenAgents puts full orchestration on top. Voice, transcription, text-based chat, and reasoning run in one vertically integrated pipeline, so responses come back in <400 milliseconds and sound human, not synthetic.

Plus, you keep full control. Plug in any LLM, integrate tools, webhooks, and MCP servers, and ground responses in your knowledge base. Get an agent live in minutes, then A/B test with Experiments, enforce Guardrails, and version every change.

The payoff: more human conversations, lower latency, and far less time stitching infrastructure together. You build on the models you already trust. Pricing is transparent and flat at $0.08 per minute.

Reply

Avatar

or to participate

Keep Reading