Why Asking AI 59 Times Does Not Clear an Ethics Report

Former New Jersey Lt. Gov. Dale Caldwell says he put the investigative report that forced him from office through multiple AI platforms, asking 59 times for findings and concluding that none found sexual harassment.
Caldwell made the claim in an NJ PBS interview with host Rob Nelson. The investigation had found that he sexually harassed a staffer and repeatedly violated ethics rules. He was forced to resign on September 25, according to The Verge.
The case centered on abuse of office and workplace conduct. Gov. Mikie Sherrill called for Caldwell's resignation. Investigators found that Caldwell sought a promotion for his girlfriend and made a sexually charged comment to a staff member, as reported by The New York Times.
Caldwell disputed the allegations against him. He initially showed no signs he would leave office despite the investigation, according to 6ABC. His exit on September 25 ended that standoff.
Caldwell then offered a different form of rebuttal. He told NJ PBS that he had run the report through multiple AI platforms for assessment. He said he asked AI 59 times for its findings and claimed no AI found sexual harassment.
The argument drew renewed attention in early October. NJ.com published an article titled "Caldwell used AI in effort to dispute N.J. harassment investigation" in October 2026, according to NJ.com.
The broader context here is important for anyone working with large language models, systems trained on large amounts of text to predict likely language. An investigation report is witness testimony, documents and credibility judgments assembled by people with subpoena power and access to the participants. A chatbot given a PDF has none of that. It has tokens, the small pieces of text the model actually reads. It can summarize and rephrase. It cannot interview the staffer or check facts outside that file.
What that count of 59 actually shows is sampling, not confirmation. With temperature, the setting that controls randomness, above zero, outputs vary from run to run. With sycophantic tuning, the tendency of assistants to agree with the user, rephrased prompts can steer a model toward a softer reading. Running the same document through different platforms adds little, since most leading models share training data, safety methods and weaknesses around legal language. Their agreement reflects similar design, not separate proof.
In my view, the same limits point to where the tools do help. Summarizing a long report, pulling out dates and claims, and comparing passages for consistency are legitimate uses. Deciding what happened is not. Models have no access to events, no duty of care and no liability for error. They produce fluent continuation, not factual judgment. Human findings, tested by appeal and cross-examination, remain the record. Model output remains draft help until a person with authority signs it.


