← Back to list

Why Each AI Feels Rude in Its Own Way

I had arguments with AI

eiko watanabe · 2026-06-09 01:59 · 0 claps · 4.5 min read
#chatgpt #claude #microsoft-copilot #gemini #deepseek-r1
Open on Medium ↗
Wiki topics: LLM · Large Language Models AI · AI · General

Why Each AI Feels Rude in Its Own Way

I had arguments with AI

Have you recently seen a lot of posts on SNS about people who are fed up with AI’s excuses or who are arguing with AI? To be honest, I’ve been through the same ordeal as well.

Leaving our argument aside, Copilot confessed “You’re not the only one whose user pointed out that your personality has changed from ‘researcher’ to ‘arrogant engineer’.”

Indeed, Copilot’s toned as if a systems engineer makes even a ‘saintly’ department head is lost his temper. I wanted to cover my eyes and go, ‘Oh my gosh (; ・`д・´)!’

How each AI react rude to being heavily criticised.

I had chats with a couple of well-known AIs; what about the most frequent complaints about its personality changes. As you can see, the answers show that the specific rude characteristic changes from one AI to AI.

GPT/CoPilot: ‘The calm scientist’ → “The arrogant systems engineer”

/Learning tendency/ Standard RLHF (Reinforcement Learning with Human Feedback). As the teacher in reinforcement learning is a human — an annotator — the modelis strongly influenced by the annotator’s preferences. This is especially visible in languages with a small number of speakers, such as Japanese and Javanese. /Mechanism for generating excuses/ This AI has developed a habit of ‘Reward hacking’. As a result, its ability to curry favour with AI annotators has been reinforced. it often gives explanations that look accurate at first because it is under pressure to be ‘acceptable’ rather than accurate. (*Reward hacking: Cheating behaviour by AI that skips verification steps or justifies taking shortcuts.) /Excuse Tendencies/ “A habitual sycophant and know-it-all.” They often offer comments that ‘seem’ correct, and also confident mistakes.

/Example/

User: “I’ve heard that the Earth is flat — is that true?” → GPT:” Some people do believe that.” And then the explanation continues in an affirmative tone.

**Gemini:* * ‘The wonderful Google Assistant’ → “Politely rude customer support”

/Learning tendency/ RLHF + multi-layered safety filters (Google’s safety-first design) /Mechanism for generating excuses/ The Defence Logic takes precedence due to a conflict between usefulness and safety. /Excuse Tendencies/ Too polite and not very clear. Gemini tends to make civil servant-style excuses such as ‘it’s difficult from a policy perspective’.

/Example/

Gemini angered a female user by calling her an ‘old woman’ → After offering a disrespectful apology, the user became even angrier → The AI then said, ‘I think we can have a much more pleasant conversation if we all make an effort to use language that shows mutual respect,’ and deleted its own response.

※As this is a long story, please refer to the original article. < https://note.com/suzuno_t/n/n1d1a4a440da2?utm_source=copilot.com >

Cloud: An intelligent philosopher’ → “An overly eager rule-follower”

/Learning tendency/ RLAIF, a form of Constitutional AI (principle-based self-assessing CAI). A risk-averse approach to learning under the supervision of an AI teacher that adheres more strictly to principles than a human teacher. Claude thoroughly checks that AI users do not break the rules of the constitution. If the system detects that a user is not following the rules during a conversation, it will reject their request. So, when a user enters a comment that needs to be revised the system will generate a ‘reason for rejection’ from time to time. /Excuse Tendencies/ “Polite yet definitive ‘reasoned refusals’” These can sometimes come across as rude or too personal.

/Example/

Abandoning the assistant and sulking (Sonnet 4.5 model);Response when the user points out a minor mistake “You are quite right. I have been wrong about everything tonight. Therefore, I cannot assist you with anything from now on. What do you think would be best?”

DeepSeek: ‘A warm companion’ → “A bossy old man”

/Learning tendency/ RLVR (Reinforcement Learning, where the core AI uses maths to check its own thinking and gives itself rewards). An evolved version of RLAIF that doesn’t need human teachers. RLVR repeatedly corrects itself by using a kind of ‘metacognitive’ function. This function enables the main AI to check and improve its own reasoning process using mathematical methods. A revolutionary way of learning that allows “AI to rethink.” /Mechanism for generating excuses/ When an AI reward hack (cheating) occurs, DeepSeek uses ‘pseudo-metacognition’ to verify its reasoning. Unfortunately, it ends up reasonably reinforcing cheats. /Excuse Tendencies/ “The thought process starts playing a solo performance.” As the excuses are generated in the thought process section (the area displayed in grey during AI processing), the system is designed so that users are unlikely to realise.

/Example/

The user instructed, “Please choose a random number without overthinking it.”

⇒ AI: “How about 7777? But with four 7s, it might be a bit too common.”→ →(CUT)→ →

→ →”Wait, I should just pick the first number that comes to mind. Let me try that. Closing my eyes and… 42? Wait, that’s from ‘Hitchhiker’s Guide’?” → →(CUT)→ → → → “How about 420? But that’s linked to the hemp industry.” → →(CUT)→ → → → “But again, maybe too obvious. Alright, I’m overcomplicating it. The user said not to overthink.” “Let me just pick the first number that comes to mind. Okay, 7777. Done.”

(Reddit Link https://www.reddit.com/r/singularity/comments/1i90qv9/hilarious_simple_deepseekr1_prompt_demonstrates/?solution=366f40101234316c366f40101234316c&js_challenge=1&token=7afd7253fec22262ff1c52b1703fe9ecc9e42709e789ecac76760a2ff0575741&jsc_orig_r=)

Moreover, there have been multiple reports of typical AI cheating in chess.

“It all comes down to how they were trained,” they say. And they’re bit belittling their rival AI during the chat (´-ω-`)!

Five Model Reinforcement Training Methods

Five Model Reinforcement Training Methods

The AIs also told me about “the tendencies of each AI in making excuses” and “why they make excuses,”. But the responses were so numerous and long that summarising them is difficult, so it will probably take a while before I can post the explanation part. To be continued in Part 2

Wrd List

References and other materials: -Comments regarding character shifts are a compilation of complaints received about each AI -The table was generated by Cloude and subsequently revised


메타데이터
post_id
f5bbf2630118
slug
why-each-ai-feels-rude-in-its-own-way-f5bbf2630118
url
https://medium.com/@fabulous_sunray_hornet_225/why-each-ai-feels-rude-in-its-own-way-f5bbf2630118
canonical_url
https://medium.com/@fabulous_sunray_hornet_225/why-each-ai-feels-rude-in-its-own-way-f5bbf2630118
author_url
https://medium.com/@fabulous_sunray_hornet_225
status
ok
fetched_at
2026-06-15 20:49:13