How to Use AI to Explore a Market You Know Nothing About

TL;DR: A client wanted deep research into the top vendors, and their offered features, in a market neither Cairo nor the customer had worked in. Three AI models, eighteen research reports, and a human making sure it gets done right - here is the workflow, and why the humans were the long pole.

A client asked us to investigate a market they weren’t competing in but were considering entering. They wanted a report on the top vendors in the space, so they could buy products and see how well those vendors actually met the market’s requirements.

The market itself was also not an area where we had experience. What we do know is how to design, build and streamline a process. So we built one, using three AI models — ChatGPT, Claude and Gemini — to develop and then verify information about those vendors.

The workflow

Flowchart of the multi-AI research process: three AI models research the market in parallel, the results are compared and ranked, each finalist vendor gets three independent deep-research reports, Claude reconciles them, NotebookLM compares the offerings and costs, a human fact-checks, and a final report goes to the customer.

The workflow as we ran it on the second study for the same client: four finalists, twelve reports. The first study, described below, carried six vendors and eighteen reports through the same steps.

We started by giving each model the same prompt and asking each for a Deep Research report. That produced three independent reports on the market.

We uploaded all three into a new notebook in Google’s NotebookLM and asked it to compare and contrast the findings, then return a single report explaining the market and ranking the vendors against the criteria we’d specified.

We read that combined report ourselves — note where the human sits in this loop — and picked the top six vendors from it.

Then we did it again, narrower. We wrote a fresh prompt for a Deep Research report on each individual vendor, and asked all three models to run it against all six. That is eighteen reports.

We didn’t stop there. For each vendor we handed the three reports to Claude, our preferred model for this kind of work, and asked it to compare them, find the places they disagreed, and go research the differences to establish which version was true. Six vendors, six comprehensive reports — each one built from three independent sources, which is the single most effective thing we found for suppressing hallucination.

Back to NotebookLM with those six, this time asking it to compare and contrast again and add a suggested bill of materials with estimated costs for equivalent hardware from each vendor. Costs came from public information: vendor websites, retail listings, and published RFQ responses.

Then humans read the whole thing again. We had taken real steps to minimize hallucination throughout, but this was a final analysis going to a customer, and both their decision and our reputation depended on the data being right. Once we were comfortable, we produced the final report. It went to the client on time, on budget, and beat what they expected.

What it cost, and what we learned

Humans reviewed at every stage, and in several places we refined the prompts or the criteria to get output that met the need. That lengthened the process. It also made the analysis better.

This was a lot of effort. Managing multiple models through multiple rounds took real time and organization — but the human reviews were the long pole, not the machines. We did it that way because we think this level of research, review and oversight is what today’s tools require.

The models produced deeper reporting in a shorter time than a person could have. Their error rate was still too high to trust implicitly. In the post-mortem our team talked about whether it had been worth the effort — it was — where the models were weak, and how to streamline it next time.

A doctor wrote in the New York Times recently in favor of using AI in medicine, and his sharpest point was that the AI doesn’t have to be perfect, it just has to be better than the doctors. This process was not perfect either, and we think the final product was more accurate than a human doing the same research alone, in less time.

It worked well enough that the client came back and asked us to run the same process on a second product category.

Recommended takeaways:

  • Use the tools. The cost is close to nothing and the product is usually better.
  • Don’t trust what they tell you. Human validation is still required.
  • Combine sources. Multiple independent reports reduce hallucination — this works with humans too.
  • Keep prompts consistent across models, or the results won’t be comparable.
  • Decide the workflow in advance. That’s what lets you shape the output to the quality your customers expect.

By the way: this was written by Mike Landry, and a human.