← Back to list

AI Society for 3.14.25 — More Deep Research

Today: More Deep Research, AI impact on language evolution, the Mercury model and dLLMs, and what’s happening to Salesforce?

dave ginsburg in AI.society · 2025-03-14 15:21 · 0 claps · 4.8 min read
#deep-research #dllm #salesforce-ai #ai-language-evolution #inception-labs
Open on Medium ↗
Wiki topics: LLM · Large Language Models CRM · Email & CRM 🔒 · Cybersecurity 🌐 · Society · General

AI Society for 3.14.25 — More Deep Research

Today: More Deep Research, AI impact on language evolution, the Mercury model and dLLMs, and what’s happening to Salesforce?

DALL-E

DALL-E

To begin, a few notes on the exploding field of AI deep research. The first by Omar Santos posting at ‘AI Security Chronicles’ is a comparison of agents that include:

OpenAI’s Deep Research

Google’s Gemini Deep Research

OpenDeepResearcher

LangChain’s Open Deep Research

Ollama Deep Researcher

He describes the differences between fully autonomous (i.e., OpenAI) and ‘human-in-the-loop’ (i.e., Google) agents and then how they operate internally including manager, tool-calling agents, web browsing capabilities and costs. The diagram from LangChain that he includes offers good background:

Source: https://github.com/langchain-ai/open_deep_research/blob/main/README.md

Source: https://github.com/langchain-ai/open_deep_research/blob/main/README.md

And performance:

· Deep research agents have made rapid progress on these benchmarks. OpenAI reported that its Deep Research agent, using the o3 model, achieved 26.6% accuracy on Humanity’s Last Exam, a dramatic leap from the ~3% that previous models like GPT-4o and Google’s Grok-2 managed. While 26.6% may sound low, this exam is so challenging that even that score far outstrips earlier AI performance, indicating a new level of expert reasoning ability.

· OpenAI and Google’s proprietary agents have the advantage of extremely powerful models (o3 and Gemini are cutting-edge, likely multimodal and trained with tool use in mind), whereas open-source ones might use weight-optimized Llama derivatives or distilled models to approximate that capability. This means proprietary agents might better handle very complex reasoning or large inputs, but open agents are quickly improving and can be run on custom hardware.

Read the full posting to for his detailed analysis, explaining why both OpenAI and Google come out on top depending upon budget and preferred approach noted above.

Next, Jim the AI Whisperer at ‘Prompt Prompts’ describes a prompt to combine ChatGPT-o3-Mini with Perplexity Deep Research to generate better essays. The beginning of his comprehensive system prompt:

Source: Jim the AI Whisperer

Source: Jim the AI Whisperer

And he notes:

For example, when I tasked Perplexity to write an article on the Titanic and its social impact on the 20th century, it took 5 minutes to complete and the article was only 1072 words long. Regular ChatGPT-o3-mini-high was faster and generated about 5000 words (although not written as well). However my combo prompt wrote 7273 high-quality words in 25 seconds.

Given that deep research generates long content, how does it impact language? Does it all begin to sound the same? Fedeminozzi at ‘Ai-Ai-OH’ asks the question, in a world where the bots generate more and more of our writing. Consider essays and articles proofed by ChatGPT, something only reinforced by long form deep research documents. And do the bots differ? Do Claude, ChatGPT, Gemini, and DeepSeek all have their own dialects? And since they are trained on internet data, do they reinforce their biases? He makes some very relevant points:

· The more we use AI to read, write, and speak for us, the more we passively accept its language choices.

· What about linguistic varieties less represented in AI training texts? These risk being excluded from the narrative. Dialects, regional variations, and languages spoken by smaller population subgroups will be underrepresented in the AI output.

· When AI writes for us, it will use the words of the majority. The standard language will overlook those who already struggle to be heard in society. This affects how we perceive these people — we can’t see or hear their words, so they cease to exist in our eyes, and their ideas lose political and social weight.

· Until now, we used language to express ideas. But if we no longer need to directly formulate them, language could shift into a tool for creating effective prompts. This may lead to a change in linguistic functions: no longer a means to interact with humans, language becomes a way to communicate with machines.

Still on models, what is ‘Mercury’ from Inception Labs, and why should we care? Ignacio de Gregorio divesinto these new models with a performance comparison and description:

The reason why these models behave so differently from AI products like ChatGPT we are used to is that they generate their responses using diffusion instead of autoregressive prediction.

And, just like image and video diffusion:

Instead of predicting one single word for every prediction the model makes, as autoregressive LLMs do, as mentioned, dLLMs predict the entire output in one go and then progressively refine it.

He concludes that dLLMs could change the game but need real-world uptake and feedback.

Then, given Salesforce’s first mover advantage for SaaS, why isn’t it a driving force in AI? Good to review the post on why SaaS and BI companies are in trouble, and if they are stuck in the ‘Innovators Dilemma.’ In any case, ‘The Information’ offers background on their Agentforce ambitions and earlier missteps with AI Cloud.

· Salesforce is pitching the artificial intelligence product as “digital labor” that can essentially replace humans for tasks such as developing sales leads and creating marketing campaigns. But many customers aren’t ready to commit to using the software, a Salesforce sales manager said

· Others say they have received incorrect answers — AI hallucinations — while testing how the software handles customer service inquiries, according to a person who does business with Salesforce. Some customers aren’t in a good position to use Agentforce because they need to connect the product to databases stored outside Salesforce to work properly.

And:

At the same time, generative AI models developed by OpenAI and other firms potentially enable companies to develop their own versions of traditional Salesforce apps at a fraction of the cost. That’s prompted questions about how much value Salesforce apps provide to customers beyond serving as a critical database for customer information.

Interestingly:

Benioff is also moving the company away from traditional seat-based software licensing to a model in which customers pay based on the computing resources they consume. That might help Salesforce avoid the negative impact from business and government pullback on IT spending.

And to close, with all the buzz around AI, where did Machine Learning (ML) go? The chart below says it all.


메타데이터
post_id
7e8965868230
slug
ai-society-for-3-14-25-more-deep-research-7e8965868230
url
https://medium.com/a-i-society/ai-society-for-3-14-25-more-deep-research-7e8965868230
canonical_url
https://medium.com/a-i-society/ai-society-for-3-14-25-more-deep-research-7e8965868230
author_url
https://medium.com/@daveginsburg
status
ok
fetched_at
2026-06-26 03:39:16