AI Testing: How to Ensure Quality in Non-Deterministic Systems with Adam Sandman
Categories: Podcasts , Test Guild
AI transforms software development by lowering entry barriers and accelerating creation, but introduces complex testing challenges in non-deterministic systems, requiring testers to adapt strategies and embrace AI tools for risk management. Quality professionals must evolve to address AI-driven complexity, advocating for their role as critical guardians of safety and compliance in an increasingly automated landscape.
Test Guild
Test Guild - hosted by Joe Colantonio has main topic focus on Testing or Automating. Each episode has a different guest. Show notes have comprehensive links and usually a full transcript. Released as audio and video.
- https://testguild.com/
- https://testguild.com/podcasts/automation/
- https://www.youtube.com/playlist?list=PL9AgRtJkydU1jqvx46esyr56BXtm1QEds
- https://www.youtube.com/@JoeColantonio
Episode Details
- Show Notes: https://testtalks.libsyn.com/ai-testing-how-to-ensure-quality-in-non-deterministic-systems-with-adam-sandman
- Published: 2026-03-10T22:33:00Z
- Duration: 43:20
- Author: Unknown
Overview
The podcast explores the transformative impact of AI on software development and quality engineering, emphasizing both opportunities and challenges. AI tools are lowering barriers to entry by enabling non-experts to build applications, accelerating development cycles, and increasing code complexity. This shift has raised the stakes for testing, as non-deterministic AI systemslike chatbots and agentic toolsproduce variable outputs requiring new testing methodologies. Quality professionals face heightened risks, particularly in critical sectors such as manufacturing and healthcare, where AI-driven failures could have severe consequences. The podcast highlights the growing need for testers to evolve their strategies, embracing AI as a supplementary tool rather than a replacement, while advocating for testing to be repositioned as a strategic business function rather than a cost center.
Key challenges include adapting to non-deterministic systems, which demand risk-based testing approaches rather than complete coverage, and addressing gaps in traditional testing frameworks. The discussion underscores the importance of decomposing applications into deterministic and non-deterministic components for targeted testing, leveraging specialized AI tools like SureWire for statistical evaluation of AI behavior. Cross-disciplinary collaboration is deemed essential, integrating risk management, data science, and domain expertise to address AI-specific challenges. The evolution of testing now also involves AI-assisted refactoring of legacy systems, compliance with regulatory standards, and ensuring alignment between requirements, code, and user expectations.
The podcast emphasizes that while AI can enhance productivity through automation and reduce manual tasks, its integration requires careful risk management, human oversight, and tailored strategies. Testers must adapt by upskilling in AI literacy, data analysis, and hybrid workflows, while organizations prioritize incremental AI adoption to address immediate pain points before tackling complex challenges. The future vision includes AI-driven development cycles, where quality assurance transitions from a purely technical role to one that aligns with business outcomes, safety, and compliance, underscoring the critical role of testers as guardians of quality in an increasingly AI-integrated landscape.
What If
-
What if you use AI to automate non-deterministic testing for a specific AI-powered feature in your product?
- Concrete move: Deploy a tool like Surewire to statistically evaluate outputs of your AI model across 10,000+ simulated user interactions, focusing on edge cases like bias detection or safety compliance.
- Why now: The text highlights that non-deterministic systems (e.g., chatbots) require new testing approaches, and AI tools can scale testing beyond human capacity.
- Expected upside: Reduced risk of critical failures (e.g., biased outputs) in high-stakes use cases (e.g., healthcare or finance), while cutting manual testing time by 60%.
-
What if you refactor a legacy codebase using AI-driven tools to modernize it and reduce technical debt?
- Concrete move: Use an AI tool like Kero to identify and remove outdated dependencies (e.g., IE 11 support) while preserving custom logic and generating unit tests for refactored code.
- Why now: Legacy systems are a bottleneck for scalability, and AI can automate modernization without requiring a large team or deep institutional knowledge.
- Expected upside: Faster deployment cycles, reduced maintenance costs from outdated tech, and alignment with modern frameworks (e.g., Entity Framework 6) for better long-term agility.
-
What if you implement risk-based testing for AI features, prioritizing high-impact scenarios over full test coverage?
- Concrete move: Partner with a risk analyst to define acceptable failure thresholds for your AI feature (e.g., 0.1% error rate in fraud detection) and use AI tools to focus testing on those scenarios.
- Why now: Teams face a 10x increase in testing demands, and risk-based approaches (from the text) allow efficient prioritization without compromising critical safety or compliance needs.
- Expected upside: Optimize testing resources by 50%, reduce false positives/negatives in critical systems, and align testing with business-critical outcomes like regulatory compliance or user trust.
Takeaway
- Adopt AI-powered testing tools like Axe or SureWire to streamline compliance testing (e.g., accessibility, regulatory checks) and reduce manual testing efforts by 50-80% for deterministic components.
- Implement risk-based testing strategies for non-deterministic AI systems by prioritizing high-impact scenarios (e.g., safety-critical workflows) and using agent-based simulation tools to evaluate variability in outputs.
- Decompose applications into deterministic and non-deterministic layers to apply targeted testing: use traditional tools for UI/data layers and specialized AI tools (e.g., agent-based testing) for AI-generated logic.
- Leverage AI for requirements alignment by using tools like Kero to cross-check code, tests, and user documentation, ensuring consistency and identifying gaps that could lead to compliance issues or functional mismatches.
- Upskill in data analysis and AI literacy by experimenting with open-source tools (e.g., Playwright, Cursor) and learning to interpret statistical outputs from AI test simulations to make informed risk management decisions.
For a PDF of longer Software Testing Podcast Episode Summaries with Briefing Notes and more detailed summary notes, visit EvilTester Patreon Podcast Summaries.