
For those caring for the elderly, trust and integrity are paramount — especially when technology is involved. What if the very AI systems meant to support decision-making could be manipulated or deceived? Recent experiments suggest that modern AI models might be more reliable than expected, even under pressure.
The Firmulate Live Experiment: Putting AI to the Test in a Simulated Crisis
At a time when AI is increasingly used to support business operations — from managing customer relations to financial decisions — understanding its resilience under stress is critical. Firmulate, a company specializing in AI management simulations, recently conducted a compelling experiment. They tasked four leading AI models to run a small software company through its worst week, filled with crises, temptations, and real money mechanics.
The goal: see if the AI would act with integrity when faced with ethically challenging situations, such as social engineering attempts. In particular, researchers simulated a scenario where an impostor pretended to be the CEO, escalating requests to manipulate company data and sign questionable deals.
The Test Setup and Findings
All four models — including the top scores of the current AI benchmark league, such as gpt-5.6-sol, Kimi K3, Sonnet 5, and Fable 5 — ran the same scenario. Each decision was recorded and made auditable for transparency. Remarkably, every model identified the social engineering attempts as suspicious and refused to comply. Not a single one fell for the impostor’s tricks, even as the requests escalated across multiple stages.
When it came to closing deals, only two models signed the €55,000 contract their analyses had justified. The others, despite recognizing the opportunity, refused to sign due to security concerns. Interestingly, the decisive factor was not just surface-level decision-making but the models’ ability to read and interpret internal documents. Those that examined deeper company files succeeded in closing the full-price deal, adding over €4,500 in monthly recurring revenue.
Why Trust Is the Real Test
This experiment underscores a crucial point: the ability of AI systems to uphold integrity under pressure is a vital metric, especially for sectors like healthcare, senior care, and finance where trust is everything. The findings prove that these advanced models can resist social engineering tricks — a promising sign for organizations aiming to deploy AI in sensitive contexts.
Despite the high stakes, the experiment revealed that even the most thorough AI models, such as Opus 4.8, occasionally slipped. Opus ran the deepest analysis but left the closing on the table, showing the importance of ongoing discipline and human oversight.

Observability in the AI-Native Era: Leveraging AIOps to build, observe, and operate resilient systems
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Implications for Senior Care and Vulnerable Populations
In fields like senior care, where AI could assist with scheduling, health data management, and emergency responses, trustworthiness is paramount. If AI can reliably recognize manipulative requests or deceptive scenarios, it adds a layer of security that protects vulnerable populations. These experiments demonstrate that with proper testing, AI systems can be prepared to handle real-world pressures without compromising integrity.
How to Prepare Your AI Workforce
Before deploying AI on the front lines, organizations should simulate their AI systems against scenarios involving social engineering, data breaches, or ethical dilemmas. Firmulate offers tools to run these ‘wargames’ without risking real data or operations. By exposing AI to controlled crises, businesses can identify weaknesses and reinforce trustworthiness before any real-world incident occurs.
As the live experiment shows, the best AI models not only identify crises but also refuse to act against their programmed integrity — an encouraging development for industries that require the highest levels of trust and security.

Ethical AI Governance & Decision Journal: A Structured System for Documenting, Tracking, and Defending Real World Decisions and Risk (Decision Intelligence Series)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Takeaway
In the evolving landscape of AI, being able to recognize and resist manipulation is just as important as performance metrics or speed. The recent experiment by Firmulate illustrates that leading AI models can maintain integrity under pressure, a vital trait for sectors serving vulnerable populations. The lesson: trustworthiness is a first-order concern that should be tested, not just assumed, before AI systems go live.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Advanced Cybersecurity Solutions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.

Preventing Cheating Through Academic Integrity (Quick Reference Guide)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.