This prospective DECIDE-AI stage 1 evaluation assessed SHAKED, a clinical decision support system constructed from multiple large language models, in a tertiary emergency department. The study aimed to measure real-world clinician adoption, temporal adoption dynamics, safety signals and preliminary clinical effects during routine ED operations over a four-week period. Authors emphasize that the evaluation was intended to inform randomized trial design rather than to support immediate clinical deployment.
The trial ran for four weeks and compared two parallel ED units in a tertiary hospital: one wing in which clinicians had access to the SHAKED system and one wing that followed routine rotations without the system. In total, 1,138 patient visits were analyzed across the two units. The report follows DECIDE-AI stage 1 principles for early prospective evaluation of AI-driven clinical interventions.
SHAKED is described as a clinical decision support system built on multiple large language models. Representative system outputs were provided in the article figures. The documentation reports that expert review sampled 100 outputs from the system and rated 99 of them as clinically appropriate. Specific internal architecture details, model names, prompt design, user interface elements and integration workflow are not fully described in the previewed text; readers are referred to the full article for comprehensive technical methods.
Clinical adoption of SHAKED by treating physicians declined over the study period, falling from an initial 68% to 30% by the end of the evaluation. Logistic analysis linked disengagement to workload: each additional shift hour was associated with lower odds of SHAKED use (odds ratio 0.72 per shift hour, 95% CI 0.62 to 0.83), indicating that clinicians were less likely to engage with the system when workload increased.
Physicians demonstrated a preference for using SHAKED in the context of radiology consultations, with higher odds of use for those consultation workflows (OR 2.98, 95% CI 1.58 to 5.63). The authors interpret these uptake patterns as evidence that clinician-perceived utility in specific tasks may influence adoption, and that operational pressures in the ED reduce sustained engagement.
No adverse events attributable to SHAKED were detected during the evaluation. Measured patient-centered operational outcomes included emergency department length of stay and consultation cycle time. Length of stay was identical between the two wings (4.9 hours in both, P = 0.99). An intention-to-treat analysis found a non-significant trend toward shorter consultation cycle time in the SHAKED wing (−9.4 minutes, P = 0.077), which did not reach conventional statistical significance.
The authors report that expert review favored the system outputs in nearly all sampled cases (99 of 100), supporting output appropriateness despite limited sustained clinician engagement.
The study's central conclusion is that sustained clinician engagement, rather than algorithmic accuracy, may be the primary barrier to effective use of AI-based clinical decision support in emergency departments. Although SHAKED outputs were judged clinically appropriate and there were no detected safety events, rapid decline in clinician use under increased workload suggests operational and human factors challenges.
Authors state these results should inform the design of randomized trials of AI CDS interventions but explicitly note that the current findings do not justify routine clinical deployment of such systems at this stage. The report therefore frames the evaluation as hypothesis-generating and preparatory for more definitive testing.
The study dataset has been released as a de-identified, patient-level analytic file accompanied by a data dictionary; direct identifiers, exact timestamps and free-text fields were removed. For legal and ethical reasons under applicable Israeli health privacy law and the study's IRB approval, the dataset is available under controlled access rather than open release. Prospective requestors must meet institutional affiliation and ethics conditions, submit a written analysis plan, and enter a data-use agreement; decisions on access are handled by the corresponding author together with the Rambam Health Care Campus Institutional Review Board.
The analysis and figure-generation code reproducing the reported quantitative results is publicly available under the MIT license on GitHub and archived on Zenodo. The repository includes the full analysis pipeline and documentation but contains no patient data. A data preparation script that reads the protected raw records is present for transparency but cannot be executed without controlled-access raw files.