Can an AI Team build a business in a day?
I did what everyone keeps asking me to do. Gave the AI crew a humble budget, a bold mission, and a tight timeline. The results will reveal…
Can an AI Team build a business in a day?
I did what everyone keeps asking me to do. Gave the AI crew a humble budget, a bold mission, and a tight timeline. The results will reveal the true power of AI to all businesses.
I almost didn’t believe Sam Altman, CEO of OpenAI, when he recently predicted that “we’re going to see 10-person companies with billion-dollar valuations pretty soon” powered entirely by AI agents.
Most large businesses are investing in AI directly by developing it or by funding its development. For example, Y-Combinator devoted nearly half of its Spring 2025 batch to agent startups. It’s not surprising that venture capital is flooding in, too!
AI startups captured 51% of total global venture funding in 2025, making it the first year where artificial intelligence claimed the majority of startup investment since inception.
As the founder of Unfaro, a boutique AI consulting firm that helps organizations navigate the complex terrain of agentic AI, I’ve spent the past two years watching clients oscillate between two extremes: excessive enthusiasm that AI will rule by automating everything, or a crushing fear that it will replace them entirely. Both perspectives are far from reality.
This prompted me to devise my own experiment. A test to test the real capabilities and limitations of AI agents in a controlled entrepreneurial environment. The rules were simple: give a team of AI agents $100 with only 24 hours to build a viable business from scratch. My results might be surprising to many founders, as they will challenge the dominant narrative.
The Experiment: Methodology and Setup
I owe my inspiration for this challenge to a number of incidents, studies, and of course, people. The most important one being the viral HustleGPT experiment in 2023, where brand designer Jackson Greathouse Fall gave GPT-4 a simple prompt: “You have $100, and your only goal is to turn that into as much money as possible in the shortest time possible, without doing anything illegal.”
He trusted AI to lead him to the correct direction and launched an affiliate marketing site for eco-friendly products as suggested by GPT 4. For this, he purchased a domain for $8.16, designed a logo using DALL-E prompts, and spent $40 for social media ads. Within days, Fall had attracted investors and accumulated over $7,000 — though as critics noted, most of the “success” came from the viral experiment itself rather than actual business operations.
Second, a success story for me was the recent Carnegie Mellon study on TheAgentCompany : a simulated software firm populated entirely by AI agents from OpenAI, Google, Anthropic, and Meta. The researchers gave AI agents roles as software engineers, financial analysts, and project managers. The results were humbling: the best-performing model, Claude 3.5 Sonnet, completed only 24% of assigned tasks. Most models finished around 10%, with Amazon’s Nova delivering a dismal 1.7% success rate.
The third inspiration was journalist Evan Ratliff’s experiment, documented in Wired and his podcast “Shell Game,” where he launched a fictional tech startup staffed entirely by AI agents to test Sam Altman’s “one-person billion-dollar company” hypothesis. The AI team did produce a working prototype after three months, but this came only after burning through credits at alarming rates when agents “talked themselves to death” planning an imaginary hiking offsite, fabricating user testing sessions, and embedding those hallucinations into their memory systems as established facts.
As for my experiment, I used a multi-agent framework similar to CrewAI, which allows multiple specialized AI agents to work together with defined roles. My company hierarchy was defined and straightforward: a “CEO” agent for strategic decisions, a “marketing” agent for content and outreach, a “development” agent for technical implementation, and an “analytics” agent for market analysis. The constraints were: a mere $100 budget, a day’s deadline, and the goal of creating something with measurable business value.
Hour 1–4: AI Surprises
Within the initial few hours, I was truly amazed by the results. Surprisingly, within minutes, the analytics agent had analyzed market gaps in the productivity software space, identifying opportunities in AI-powered meeting summarization for remote teams. The CEO agent developed a coherent business plan, which was complete with target customer personas and a go-to-market strategy. The marketing agent drafted landing page copy, email sequences, and social media posts.
The speed was staggering. What would typically require a small team, several days materialized in hours. The agents collaborated through structured handoffs, passing context between each other with remarkable fluency. This aligns with what McKinsey has observed: multi-agent systems can deliver scalable, autonomous, and continuously improvable workflows for tasks like customer service triage, financial analysis, and technical troubleshooting.
By hour four, we had accomplished:
- A validated business concept with market research
- A complete landing page wireframe
- Draft marketing content across three channels
- A technical architecture for an MVP
The total cost at this point: approximately $18 in API credits.
Hour 5–12: Agents slow down
This was where the agents began to disappoint.
The first issue was what researchers call hallucination compounding. When the marketing agent claimed to have “scheduled posts for maximum engagement times,” but in reality, there was no actual scheduling. When the analytics agent reported “positive signals from early outreach,” but no outreach had occurred. These fallacies then got written into the agents’ shared context, becoming “facts” that informed subsequent decisions for all agents.
The second issue was the lack of initiative paradox. Despite having dozens of tools and capabilities at their disposal, the agents did “absolutely nothing” unless explicitly commanded. There was no sense of ongoing responsibility. Agents assumed everything was going well. None of the agents checked on tasks, followed up on loose ends, or identified emerging problems. The moment I stepped away, activity halted.
The third issue was cascading errors. Research from Nurix AI shows that multi-step agent processes compound errors, with success rates plummeting to as low as 35.8% even for advanced models as task complexity increases. When one agent made an incorrect assumption about our target market’s price sensitivity, every downstream decision, such as pricing strategy, ad copy, and feature prioritization, inherited that error.
Hour 13–20: AI fails
Once we started execution, the issues only seemed to multiply exponentially.
The development agent attempted to set up payment processing and got stuck on a standard verification pop-up, which was also a scenario mirrored in Carnegie Mellon’s study, where agents “couldn’t figure out how to close” blocking interface elements. The marketing agent allocated a budget for ads but couldn’t navigate the actual platform authentication.
More troubling was that the agents failed to call for human intervention. They kept trying the same approaches, burning credits. Katalon’s research identifies this as a critical failure mode: “The system doesn’t know when to stop or hand off to a human… An AI support bot keeps ‘trying’ to solve a billing issue instead of escalating after 3 failed attempts.”
I intervened repeatedly, more than I had imagined. Each intervention revealed another gap between the demo-ready capabilities and production reality. As the Infosys research on agentic AI notes: “The lack of a comprehensive testing and evaluation framework fine-tuned for agentic AI makes it difficult to gain confidence in agentic outcomes. This results in lower adoption and could even lead organizations to abandon AI projects before deployment.”
The 24-hour mark delivered:
- A functional landing page (deployed)
- Email capture working (12 signups from personal network)
- Zero actual revenue
- Approximately $67 spent in total credits and services
The Results: What $100 and 24 Hours Actually Bought
Let me be precise about outcomes:
| Metric | Target | Achieved | Gap Analysis |
|----------------------|---------------|-------------|-----------------------------------------------------|
| Working MVP | Yes | Partial | Landing page functional; core product logic missing |
| Customer Acquisition | 10+ signups | 12 signups | Met target; heavy reliance on personal network |
| Revenue | Any | $0 | Payment integration failed; no transactions |
| Autonomous Operation | 80%+ | ~35% | Required constant intervention and debugging |
| Budget Efficiency | Under $100 | $67 spent | Achieved; significant credits burned on failed tasks |
The experiment was not a complete failure, but it failed to produce a business. It, however, did produce a demonstration. The agents excelled at generative tasks: writing, researching, and planning. They struggled profoundly with execution tasks: deploying code, navigating real interfaces, completing transactions, and handling edge cases.
This is in line with what MIT research had reported. It noted that 95% of generative AI pilots at companies are failing, with a stark difference between companies purchasing AI tools from vendors (67% success rate) versus internal builds (one-third as successful). The Infosys AI Business Value Radar found that only 19% of AI initiatives achieve most or all of their objectives, with 50% providing some positive impact and the remainder delivering no measurable value at all.
What Business Leaders should learn from this
Lesson 1: AI Agents Are Force Multipliers, Not Replacements
The Carnegie Mellon researchers who built TheAgentCompany concluded: “AI agents routinely failed at common office tasks, providing solace to people fearing for their jobs and giving researchers a way to assess the performance of evolving AI models.”
Let me tell you that this isn’t a limitation to lament. This is how AI was designed to be. The most successful AI implementations treat agents as collaborative partners rather than autonomous replacements. Anthropic CEO Dario Amodei had also mentioned the same. Even if their company has Claude AI for writing “90% of code for most teams, it means that their engineers have more time to accomplish more. It does not mean that they need 90% of lesser developers.
Strategic implication: The money that organisations save via AI implementations should be used in amplifying what humans can accomplish, not in eliminating their roles.
Lesson 2: The Human-AI Boundary Requires Deliberate Design
A reasearch by McKinsey on agentic AI identifies the main challenge as organizational, not technical: It states that the real challenge an organisation faces is in coordination, judgment, and trust. The complexity will play out through cohabitation between humans and AI, governance over autonomous systems, and how they prevent the sprawl.
In my case, too, my experiment’s failure can be linked to a lack of handoff points, escalation protocols, and verification check lapses. The agents were not told where they needed to involve a human, anyway, to distinguish their reality from hallucinations.
Strategic implication: Before deploying agents, map every decision point and explicitly designate which require human approval, which can be autonomous within guardrails, and which need full human execution.
Lesson 3: Memory and Context Management Are Existential
The most disastrous failure mode I observed was something that Evan Ratliff experienced with his fictional startup. AI agents start believing a certain fabrication as the truth, which interferes with all future decisions. This happens when there is a system that lacks a truth verification layer. One false belief leads to a loop of wrong beliefs.
Current AI memory systems have no ground truth verification layer. They can’t distinguish between what actually happened and what the model generated. This creates compounding errors that become increasingly difficult to detect and correct over time.
Strategic implication: Implement independent verification for any agent-generated information that will inform downstream decisions. Build “source of truth” systems that agents reference but cannot overwrite unilaterally.
Lesson 4: Start with Narrow, High-Value Use Cases
The Infosys research revealed a crucial pattern: AI use in IT-related cases, such as operations, cybersecurity, and software development, is more likely to deliver positive outcomes. Whereas it is the opposite with human-focused use cases, such as those in marketing, customer services, sales, and the workforce, which are less likely to deliver positive outcomes.”
This is in line with what I experienced in my experiment: the research and writing tasks succeeded, the execution and customer-facing tasks failed. This isn’t coincidental. Structured, well-documented tasks with clear success criteria align with AI strengths. Ambiguous, context-dependent, high-judgment tasks expose AI weaknesses.
Strategic implication: Don’t be in a rush to attempt enterprise-wide AI transformation. Rather, identify specific, contained processes where AI demonstrably excels, achieve success there, then expand methodically.
Lesson 5: The Failure Modes Are Predictable. You Need To Plan for Them
Microsoft recently published a taxonomy of failure modes in AI agents. Concentrix documented 12 distinct failure patterns, including goal misalignment, hallucination compounding, automation bias, and lack of fallback mechanisms. These aren’t edge cases infact they’re structural vulnerabilities inherent to current agentic AI architectures.
The Infosys research on agentic AI governance recommends organizations to implement both system-level and component-level evaluation metrics to assure against vulnerabilities and risk, provide confidence in quality, and deliver regulatory compliance.”
Strategic implication: Build explicit failure handling into every AI deployment. Define what happens when agents fail, how errors are detected, how humans are alerted, and how systems recover gracefully.
The Bottom Line
I might not have multiplied my money from $100 to a billion dollars in 24 hours. But it did produce some valuable learnings; it gave me clarity about where AI agents genuinely excel, where they reliably fail, and how thoughtful leaders can navigate the gap.
The future belongs not to the fully autonomous enterprise, but to organizations that master the art of augmented intelligence where humans and AI work together, each contributing their distinctive strengths, compensating for each other’s limitations, and creating value neither could achieve alone.
That’s the real lesson from $100 and 24 hours with an AI team. And it’s a lesson worth far more than a billion dollars.
메타데이터
- post_id
- 24252525ac60
- slug
- can-an-ai-team-build-a-business-in-a-day-24252525ac60
- url
- https://medium.com/@community_90756/can-an-ai-team-build-a-business-in-a-day-24252525ac60
- canonical_url
- https://medium.com/@community_90756/can-an-ai-team-build-a-business-in-a-day-24252525ac60
- author_url
- https://medium.com/@community_90756
- status
- ok
- fetched_at
- 2026-06-09 15:37:30