We've all seen people talking to ChatGPT or Grok on their phones asking very simple questions, queries known in the AI industry as prompts. Some of them even think they use it better than most because they have seen these kinds of posts on social media:
Bad: "Give me 20 marketing ideas for my business."
Good: "Give me 20 creative marketing ideas for my small business that will help me stand out from competitors."
Great: "Think like a world-class CMO in [Blank] industry. Give me 30 highly creative, out-of-the-box marketing ideas for my business. Make sure they address my target audience's biggest pain points."
Three boxes: bad prompt, good prompt, great prompt. The "before" looks flat and the "after" looks sharp. It's one of the most shared formats on social platforms, and most of that advice is only teaching people to fail more elaborately.
Read those three again, in order. Twenty ideas became thirty. "Creative" became "highly creative, out-of-the-box." A world-class "Chief Marketing Officer" got invoked. And in all three versions, nothing is properly defined. Nobody says what the business actually sells, who buys it, or what those buyers are actually struggling with. Not once.
So, the AI model does the only thing available to it: it invents a plausible audience, assigns that invented audience a plausible pain point, and writes thirty confident ideas for a business that could be anything and a customer who doesn't exist. The "great" version looks more finished. It looks more credible. But the information stayed at zero. So, it looks applicable to the user, as would anything so generic.
Here's what actually happens next, and it's the real cost of this whole genre of advice. Someone runs the "great" version, gets thirty generic ideas that could work equally well for a dental practice, a taquería, or a law firm.
So, what does a single prompt that actually works carry? Here is one example:
Do you notice what's doing the work? It isn't length. It's five specific things.
First, it names a specific specialization, a job description instead of a job title — "valuing mixed-use and retail properties in Mexico City's central corridors." Not "a real estate expert." A specialization that narrow gets you someone who already knows "preventa" doesn't price like titled inventory.
Second, it hands over the vocabulary the model can't be relied on to reach for: "uso de suelo", "preventa", closing price versus asking price. These are the words a specialist in this exact market actually uses. Supply them, and the model uses them correctly. Withhold them, and it either avoids the topic or fakes fluency — and you can't tell which from the output alone.
Third, it names real sources — CBRE México, JLL México, Softec, AMPI, the Sociedad Hipotecaria Federal's index — instead of leaving "research this" to mean whatever the model finds first.
Fourth, it draws a scope limit: "Do not recommend a listing price yet." One task, not five, and no premature conclusion dressed up as an answer.
Fifth, it asks for honesty. It's the highest-leverage line in the whole prompt. Whether about the user's assumptions or about the AI's findings. And this is the one most people skip entirely: to tell the model to be honest. Most prompts reward confidence by default, because confident-sounding answers feel finished. This one asks for the opposite, on purpose.
Here's a second variable most people don't think about: which AI model to use. Same prompt, different models, different quality, and most people couldn't tell you which one they should be running. For instance, most Claude users don't know the difference between Sonnet 5, Opus 5, Fable 5, and the previous-generation Opus 4.8, or what effort setting to run any of them at.
I had a PowerPoint built the way our brand materials are built, and I wanted to show someone, in real time, the practical gap between AI tools. I gave identical instructions to two AI tools on their free tiers and to one on a paid plan running its top model. The free tiers don't give you the best models. They route you to the lightweight one. Both free versions produced the same failure, overlapping, unreadable text. The paid one produced a clean result on the first pass. Same instructions, same file, different engine behind the button.
Defining which model, which tier, and which settings, web search on or off, how much reasoning effort to spend, which files the model can actually see is essential for the AI to produce a reliable output.
There is, however, a ceiling above everything that has been described so far, and no amount of prompt craft gets you past it. Even the real estate prompt above is a single question, asked once, in one conversation. And a single conversation has hard edges: how many files you can attach, how many images, how much text before the model starts losing what you told it at the beginning. That's fine for a one-off question.
However, it's entirely the wrong tool for anything that runs for weeks with documents piling up, a financial analysis, market study, ongoing negotiation, project development, etc. That's not a prompting problem anymore. It's a structural one. And the fix isn't a better one-time prompt. It's what's known as an "Agent".
An agent can be a strategist (a chat inside a folder), or it can work while you sleep — wired through an API, running on a trigger instead of on a request. They are structured with the following six sections:
- Role
- Defines who it is — the specific expertise, the same move as the real estate prompt above, made permanent instead of re-typed.
- Task
- Defines the one thing it produces, for whom, and to what standard, so output doesn't drift depending on how the request happened to be phrased that day.
- Specifics
- Carries the rules — and the reasoning behind each one, because a rule with a "why" attached gets applied correctly to situations nobody wrote down in advance. A rule with no reason gets applied literally, and wrongly, the first time it meets an exception.
- Context
- Gives it the business it's operating inside: the client, where its output goes next, what a weak answer actually costs the person downstream.
- Examples
- Is the highest-leverage of the six. These models imitate demonstrated quality far more reliably than they follow a written description of it. One excellent example generalizes better than a paragraph of instruction ever will.
- Notes
- Carries the guardrails — format, language, and what it does when it doesn't know something. The same instinct as "give me a range, not a confident number" from the real estate prompt, except built in permanently instead of requested each time.
Put together, that's not a toy. It's a specialist who is fully briefed — no drift, no re-explaining, no retraining.
That's the shape of what's actually running, right now, inside our business: specific customized agents, each holding its own files according to our clients' situations. Not a tool answering a question once. It's a system trusted to deliver over time. A specialist that can be cross-examined, conversed with and consulted back and forth for the utmost efficiency because it knows who you are and what your situation is.
The Deloitte example.
In July 2025, Deloitte Australia delivered a 237-page independent assurance review to the federal Department of Employment and Workplace Relations.
The contract was worth A$440,000. A researcher at the University of Sydney read it and found references to academic works that did not exist, including a book attributed to a professor of constitutional law on a subject outside her field. He also found a quotation attributed to a Federal Court judge. The judge never said it.
Three decisions produced that outcome, and a machine made none of them.
The first was choosing the wrong instrument, a generative AI toolchain, Azure OpenAI GPT-4o. Deloitte's own disclosure says the tool was used to address traceability and documentation gaps — that is, to produce sources. A generative model composes text that resembles a citation. It does not retrieve one. Asking it to supply references is asking a novelist for a car manual.
Deloitte used the wrong tool.
The second was the absence of a verification step. Nobody in the chain checked whether the cited works existed. An outside academic caught it in a single reading. The failure was not that the model produced an unverified claim, it was not understanding the purpose of the tool they were using.
The third was silence. The original report disclosed nothing, so there was nothing to be fixed. The disclosure appeared only after the errors were made public.
Senator Deborah O'Neill called it a "human intelligence problem." The researcher who exposed it objected not to the technology but to what he called a "non-expert methodology."
There was no exploit, no manipulation, no ill-intended user. The tool behaved exactly as configured. It was pointed at the wrong task, by people who did not know it was the wrong task, and no one downstream was checking.
The instructive part is what Deloitte did next. It did not stop using AI. In the same period, the firm announced a deployment of Claude across nearly 500,000 employees worldwide. The remedy for badly configured AI is not less AI. It is someone who knows how to configure it.
Which brings us back to where we started. Knowing what goes inside a prompt, knowing which model runs it, knowing where the single prompt ends and the agent begins is what removes the blindness and makes AI a valuable tool for businesses.