My Research Agent Worked. So Why Wasn’t I Using It?

Key Takeaways

  • An agent can complete its assignment and still produce something you rarely use.

  • Checking sources, reviewing samples, and noticing recurring problems are part of supervising an agent.

  • The value of the result needs to justify the time spent reviewing and correcting it.

In my last post, I shared how I worked with an AI agent to build a school psychology video game. Much of that process involved playing the game, noticing what wasn’t working, and revising it together.

In an earlier post, I used my research system to explain how goals, repeated steps, quality checks, stopping rules, and handoffs work together. On paper, the assignment seemed clear: preserve new articles in a searchable archive and return a short digest of the items most relevant to my work.

Then I had to live with the system. That exposed a different question: What happens when an agent completes its assignment, but the result still is not useful enough—or requires more supervision than expected?

I receive a steady stream of research alerts related to artificial intelligence, school psychology, psychological assessment, training, and mental health. Staying current supports my research, teaching, consulting, and writing. But the incoming information had become a fire hose.

The first version of my Research Library agent did real work. It could identify articles, remove duplicates, assign priorities, and place the results in a searchable library. The collection was far more organized than the emails it replaced.

There was only one problem: I was not consistently using what it produced.

An Organized Fire Hose Is Still a Fire Hose

My original goal was essentially to collect and organize new research. That sounded sensible because the incoming alerts were the visible problem.

But even with priorities and categories, I was being handed more information than I could reasonably review. The agent had reduced the work of collecting research while leaving too much of the work of deciding what deserved my attention.

Imagine asking a capable intern to review a stack of articles. The intern returns with every article carefully labeled and placed in a filing cabinet. The assignment was completed, but you still have to open the cabinet and decide what matters.

A more useful result would include a short briefing: here are a few articles worth your attention, why they may matter, and where to find them.

That was the part of the assignment I had failed to design.

Keep the Archive and Shrink the Digest

The revised goal was to help me notice a small number of useful articles while retaining the full collection for later.

The archive and the digest serve different purposes. The archive stays searchable because an article that seems unimportant today may become useful for a future study, class, or presentation. The digest brings forward a limited selection so I can scan it without reopening the entire library.

Each entry needs enough context to help me decide whether to read further: a brief summary based on a verified abstract when available, why the article may matter, the country or setting and sample when reported, and direct links to the source and library record.

Those details also make the work easier to check. A study’s setting and sample can affect how I interpret its relevance. A source link lets me compare the summary with the original. If the agent cannot verify a detail, it should leave the uncertainty visible.

This is still a work in progress. For example, an entry may say, “Study country: Not stated in the abstract.” In my experience, this often happens with U.S. studies whose authors do not identify the country in the abstract. That does not establish where a particular study took place; it shows a limit of looking only at the abstract.

I still need to update the assignment so the agent searches the full article, when accessible, for evidence about the study location. If the full text is unavailable or the location remains unclear, the entry should say so. The fix is to look for more evidence, not to treat an omitted country as proof that the study was conducted in the United States.

The design lesson extends beyond research: preserve what may matter later, but make the immediate result manageable for the person receiving it.

The Problems Made the Assignment Clearer

Trying to use the library exposed weaknesses that an organized collection alone could hide.

Article retrieval was one. A promising link might lead to an abstract, a login screen, or an error page. Opening the page did not mean the agent had obtained the article. The saved file needed to be checked before a download could be called successful. Repeated attempts at the same blocked route added effort without solving the problem, so limits and clear records of failed attempts became part of the design.

Audio was another experiment. I had read blog posts and listened to podcasts in which others described using agents to pass articles to Notebook and create audio overviews. I wanted to listen to research while walking or driving, so I had my agent hand articles off to Notebook for that purpose. The agent handled the handoff; Notebook was the tool intended to generate the audio.

In my setup, getting from selected articles to completed, playable overviews took more supervision and troubleshooting than it was worth. Having an article queued did not mean an overview had been generated. Someone else’s successful demonstration did not translate into an equally useful process for me.

Your mileage may vary. The tools, permissions, sources, and time you can spend troubleshooting all affect the result. I decided to make the audio work a separate scheduled task that runs independently on a regular daily schedule. That keeps it separate from the main research digest, so a problem creating an overview does not have to hold up the written material. Separating the tasks does not remove the need to check whether the audio was actually produced.

Feedback also needed more than an appealing interface. An early browser-based feedback feature did not produce the dependable saved record needed to inform later priorities. It was retired. The lesson was straightforward: the agent should not imply that it is learning from my feedback unless that feedback is actually preserved and used.

These were lessons about the assignment itself. I had to specify what counted as a completed result, what evidence should accompany it, and when the agent should stop.

Check Whether It Is Still Helping

Improving the initial design does not end the supervision. The question becomes whether the agent continues to produce work that is accurate, useful, and worth the attention it requires.

For the research agent, a practical review should include a sample of articles it labeled successfully processed. Do the links lead to the right sources? Do summaries match the abstracts? Are important qualifications preserved? Do the priorities make sense for my work?

Reviewing only the items flagged as failures would miss mistakes in polished, apparently successful entries. Review should be closer when the agent is new or the assignment changes. A history of good results can support cautious confidence, but it does not eliminate the need to check.

The agent also needs to leave an understandable record of what it examined, added, skipped, or could not verify. That makes it possible to distinguish one bad link from a recurring retrieval problem, or an occasional correction from a pattern of inaccurate summaries.

This is the same lesson the game made visible. I had to play it to see whether the experience made sense. With the research agent, I need to read and use the results to judge whether the assignment is being accomplished.

Know When to Change the Assignment

A useful agent can become less useful as circumstances change. A source website may move its pages, an application may change, or my research interests may shift. Those changes are reasons to recheck the agent’s work and access.

Repeated errors should prompt a response. Depending on the problem, that could mean clarifying an instruction, narrowing the assignment, changing a source, adding a check, reducing access, or pausing the agent while the problem is examined. Repeating the same correction indefinitely is not a good supervision plan.

The time required for supervision matters too. An agent that saves an hour of searching but creates two hours of cleanup is not saving time. It may still offer another benefit, such as a more searchable record, but that benefit needs to justify the effort.

The practical questions are simple: Am I using this? Can I check it? Are the same problems recurring? Is it still worth doing this way?

Sometimes the answer will be to make the agent’s job smaller. Sometimes ordinary software or completing the task directly will make more sense.

Bringing the Series Back to Practice

The research library and the game gave me different ways to see the same issue: giving an agent a goal is only the beginning. The result has to work for the person who will use it, and someone has to remain responsible for checking that it does.

Research monitoring can be a useful starting point because it can use public professional information without involving student or clinical records. Even here, an agent’s summary does not replace reading a source or exercising professional judgment about its relevance.

The intern analogy still fits. A capable but uneven assistant needs a clear assignment, examples of useful work, access appropriate to the task, and a supervisor who can inspect and reject the result. Reliable past performance does not transfer that responsibility to the assistant.

That is where I want to leave the main series: start with a task you understand, make the expected result concrete, try what the agent produces, and use what you learn to improve the assignment.

We Are Still in the Early Days

We are still in the early days of making agents dependable parts of everyday professional work. These projects have given me useful results, but they have also required revision, failed experiments, and decisions to simplify. A polished demonstration can show what is possible without showing all the effort needed to keep it working.

Your mileage may vary—and you may not want to get into the car yet. That is a reasonable choice. You can learn what these tools do and follow their development without giving an agent access to your files or adding another system to an already full workday.

I see agents as part of the future of how we work—and, for some of us, already part of the present. My research library and the game are examples of that shift in my own work. That does not mean every task needs an agent, or that everyone needs to adopt one now. It does mean I think the ability to define an assignment, set limits, and judge the result will become increasingly useful.

Final Takeaways

  • Define what you need to receive before building a long list of steps for the agent.

  • Keep sources and missing information visible so the result can be checked.

  • Review some apparently successful work and respond to recurring problems.

  • Include your review and correction time when deciding whether to continue, simplify, or stop.

I have also posted a supplemental delegation audit and calculator to help you apply these ideas and decide which tasks—or parts of tasks—might be appropriate to hand to an agent. Read the delegation post here: __________.

Earlier Posts in This Series

AI Disclosure: Generative AI was used to assist with drafting and editing this post and to create the accompanying image. I reviewed and revised the content and take responsibility for its accuracy and final form.

 
Adam Lockwood

Adam B. Lockwood, PhD, NCSP, LP, is a school psychologist, researcher, and consultant focused on the responsible use of artificial intelligence in education and psychology.

https://lockwoodconsulting.net/about
Next
Next

I Built a School Psychology Video Game with an AI Agent