After I was laid off, I kept circling an itch I could not quite name.
AI was becoming capable of doing more kinds of work, quickly. But I could not see clearly what had to exist between a model doing something impressive and a business being able to rely on it. What made an AI-supported decision trustworthy enough to act on? How did a system know whether its information was current, authoritative, complete, or connected to what was actually happening?
I did not begin with a product thesis or a claim that I had found a new problem. I had a pattern I wanted to understand.
So I started exploring it with ChatGPT. We followed questions about evidence, current state, provenance, authority, exceptions, and what happens when an AI makes a perfectly reasonable decision from an incomplete picture of reality.
The deeper we looked, the more of that territory turned out to be established work with existing names: data governance, workflow systems, business rules, process management, human factors, and decision support. The possible business ideas that emerged along the way got smaller, more crowded, or less credible.
The research was doing its job.
But it eventually raised a different question: if the same AI was helping me generate the questions, find the prior art, attack the ideas, interpret what survived, and identify the next thing to explore, how independent was the process?
I started treating each new idea as something that needed to survive research, not just make sense in the conversation.
One of the rules I adopted early was simple: new to me did not mean new. When something looked promising, we went looking for the people who had already studied it, the products already addressing it, and the terminology I did not know yet.
I’ve already used “we” a few times, and even that compresses what was actually happening: I was directing the inquiry while ChatGPT searched, synthesized, assessed and responded inside the same conversation.
Provenance was one example. Knowing where information came from, whether it was current, and what had authority over it seemed important enough that I wondered whether there was a business opportunity there. The research pointed instead toward decades of existing work in data governance, business rules and workflow systems, along with newer products already addressing the problem. What looked like a possible gap became a much narrower question about how established practices change when generative AI becomes part of the decision path.
Other paths ended differently. Some were real problems but poor business opportunities. Existing platforms were already absorbing capabilities I was imagining, and other areas were crowded with companies that knew the territory better than I did.
So the questions kept changing.
Was the problem real? Was it already understood? Had AI changed it enough to matter differently now? Was there something useful to build, or was I simply learning an unfamiliar field?
Eventually, the larger business idea stopped surviving them.
That mattered because the process was giving me more than reinforcement.
ChatGPT could find prior art and context I did not know about, identify counterarguments I had not considered, and then make its own assessment of what that evidence meant for the idea we were examining. Sometimes that assessment held up when I looked at the underlying evidence. Sometimes what we found was enough for me to abandon an idea.
The process had done a lot of what I wanted it to do. I started with questions I did not know how to frame, learned the language of fields I had never worked in, found research I would not have known to look for, and abandoned ideas as I learned more about the territory.
My claims were getting narrower. I knew more of the prior art. I was getting better at separating an interesting problem from a business opportunity, and more explicit about where my own knowledge stopped.
Those were concrete signs of progress. They were also part of why I trusted the process.
What I had not asked was what, exactly, they gave me reason to trust.
ChatGPT could argue against an idea it had helped me develop. It could find evidence that weakened its earlier assessment. It could tell me that a business case no longer held up and explain why.
Those were useful checks. But they were still happening inside the same collaboration.
I had made the process more adversarial. I had not necessarily made it more independent.
Nothing had obviously failed. In many ways, the process had worked. What I could no longer tell was which kinds of mistakes it was well positioned to catch, and which might remain invisible inside the collaboration.
At one point, I put the problem to ChatGPT more directly: were we actually finding the gap, or were we just performatively iterating?
The question had been bothering me because iteration was clearly happening. We were finding evidence, changing the questions, narrowing claims and abandoning ideas. I could open the sources ChatGPT found and decide whether they supported what it was telling me.
But that only answered part of the problem.
Checking a source gave me another point of contact with the evidence. It did not make the inquiry itself independent. ChatGPT was still helping decide what questions were worth pursuing, finding much of the evidence I encountered, interpreting what it meant, challenging the conclusions we had reached, and helping determine where to go next.
I was still making the decisions. That mattered. But being the person with final authority did not mean I was independently validating everything that informed those decisions.
In areas where I lacked expertise, that distinction became harder to ignore. I could challenge an answer, ask for contrary evidence, inspect sources, or reject a conclusion. But I was often relying on the same system to help me understand what I should challenge, what evidence might matter, and what I might be missing.
One very useful collaborator had begun occupying several roles in the process at once.
I was the human in the loop. I chose what to pursue, decided when an idea had failed, and could reject anything ChatGPT produced.
But authority and independent judgment are not the same thing.
There were parts of this work where ChatGPT had capabilities I could not practically match. It could search across unfamiliar fields, surface terminology I did not know, synthesize large amounts of material quickly, and help me reason through technical arguments outside my expertise.
That was the point of using it.
It also meant that my judgment sometimes depended on work I could evaluate only partially on my own. Telling me that I remained responsible for the final decision did not change that.
My experience here also has a boundary. I was working largely as an individual, not as a subject-matter expert inside an enterprise where architecture reviews, domain specialists, security teams, legal review, testing, or other structures might introduce additional expertise, evidence, or challenge.
Those structures do not guarantee better judgment. What matters here is that my process did not naturally contain them.
What I was experiencing is closer to something increasingly visible in independent development and vibe coding: AI allows one person to work across expertise boundaries that would otherwise require different specialists, tools, or considerably more time. The person can remain in charge while relying on capabilities they could not independently reproduce or fully evaluate.
The same pattern can extend beyond software.
People have always made decisions using evidence, analysis and expertise supplied by other people. We trust researchers, analysts, attorneys, engineers and colleagues to know things we do not.
But those roles have often been distributed across different people, tools and moments in time. The person proposing an idea may not be the person researching it. The person challenging it may bring different expertise or incentives. Evidence may come from somewhere neither person controls.
That separation does not guarantee good judgment. People share assumptions, repeat one another's mistakes and reinforce bad ideas all the time.
AI changes the shape of the problem because so many of those roles can now collapse into the same interaction.
What I needed to understand was whether the loop itself contained enough separation for my judgment to mean what I thought it meant.
Seeing the problem did not make me want to use AI less.
It made me more deliberate about what I was asking it to do.
The useful distinction for me became less about which tasks belonged to the human and which belonged to the AI, and more about where the process needed another basis for judgment.
Sometimes that means creating separation from the collaboration, but separation by itself is not the point. What matters is whether the check has a meaningful chance of exposing something the existing process could miss.
That can take different forms. I can leave the conversation and inspect the underlying source. I can test a claim against something observable in the real world. I can ask someone with relevant expertise to challenge an assumption I am not qualified to evaluate myself.
Another model can still be useful. Giving the work to a model that did not participate in developing it can reduce some of the influence of the original conversational history and expose assumptions that became invisible inside it.
That is contextual distance, not independent verification. The new model may share many of the same limitations, assumptions, or sources as the first. When the question can be checked against a primary source, relevant expertise, a test, or an observable result, those provide a different basis for judgment.
The distinction matters because these checks are doing different jobs. A fresh model may help me see the work differently. It does not, simply by being separate, tell me whether the work is right.
The goal is not to make every step independent. That would give up much of what makes the collaboration useful.
It is to notice when one loop has begun doing too many epistemic jobs for itself, and deliberately introduce enough separation for the important claims to encounter something outside it.
I still want the speed. I still want the breadth. I still want to work on problems I could not practically explore this way on my own.
I just no longer think skepticism inside the collaboration is always enough.
I am still working out what enough separation looks like in practice.
It will not be the same for every kind of work. A claim in an article, a piece of software running locally, and a decision that affects customers or employees do not carry the same consequences. The amount of independent challenge worth introducing should reflect that.
What has changed for me is the question I ask.
I used to focus more on whether I was using AI critically: Was I checking its work? Asking for evidence? Challenging assumptions? Making the final decision myself?
Those things still matter. Now I also ask where the challenge is coming from.
If the same collaboration generated the idea, found the evidence, interpreted it, attacked the conclusion and helped decide what survived, I want to know what in the process has had a meaningful chance to disagree from somewhere else.
Sometimes that may be another person. Sometimes it may be direct measurement, an authoritative source, a test against the real world, or a deliberately separate review with different context.
None of those makes the process automatically correct. Another person can be wrong. A second model can fail in the same way as the first. An authoritative source can be outdated or applied badly. A test can measure the wrong thing.
The point is not to manufacture disagreement.
It is to give important ideas a chance to encounter something that did not help create them.
That is a different standard from simply keeping a human in the loop, and I am still learning when it matters enough to insist on it.
I started this work trying to understand where my experience still mattered as AI changed what I could do.
My experience matters. So does knowing where it stops.
AI lets me work across boundaries that would previously have required more time, more people, or expertise I did not have. I do not want to give that up. Learning to use that capability well is part of how I intend to keep growing as a technology professional.
But getting better at collaborating with AI cannot mean only getting better at extracting useful work from it.
It also means becoming more aware of what I know well enough to judge, what I am relying on the collaboration to help me understand, and when the difference matters.
I expect that line to keep moving. I will learn more. The models will become more capable. The kinds of work we can do together will change.
So I do not think the answer is to draw a permanent boundary between what the human does and what the AI does.
The harder responsibility is recognizing when the collaboration has carried me somewhere my own judgment cannot fully follow, and knowing when that work needs another point of contact with reality.
That feels less like a solution than a professional discipline I am only beginning to understand.
Afterword: Testing the Article
The article you just read is not the article I originally wrote.
An early draft went to someone I trust for feedback. His criticism exposed weaknesses I had not seen, including one that eventually changed the argument itself: I was thinking carefully about challenging AI output, but not carefully enough about what it meant to rely on AI to help me cross expertise gaps I could not fully evaluate on my own.
The rewrite that followed became substantial enough that I started treating the article as another test of the idea I was writing about.
I gave versions of the argument to people and models that had not participated in developing it. I wanted the work to encounter criticism without carrying along the history of the conversation that had made the argument persuasive to me in the first place.
Some of that criticism changed the article. Some did not survive when I looked more closely at what it was claiming.
One experiment was especially useful. I tried making a separate model more explicitly adversarial, expecting stronger criticism. It certainly became more skeptical. The criticism did not necessarily become better.
That was a useful reminder that critical-sounding is not the same as independent, and independent is not the same as correct.
None of these reviews verified the argument in this article. Agreement between reviewers would not have done that either.
What they did was give the article a chance to encounter something that had not helped create it.
The article changed because of what survived.
✦