Comparing Two Patient Intake Approaches for a Conversational Triage Chatbot: A Structured Reason For Visit - Type Intake Versus Initial Assessment Questions

Authors

Laurie O’Bryan, BSN, Nurse Editor, Schmitt-Thompson Clinical Content

David A. Thompson MD, Senior Medical Editor, Schmitt-Thompson Clinical Content; Northwestern University Feinberg School of Medicine

Karen Tancock RN BScN, Digital Health Analyst, Schmitt-Thompson Clinical Content

Background: Conversational triage chatbots using large language models (LLMs) and retrieval-augmented generation (RAG) show promise for automating telehealth nurse triage. A key design question is how best to gather patient information before the chatbot applies clinical triage logic. We compared two intake approaches. The first, Initial Assessment Question (IAQ) Intake, uses standardized assessment questions found in each telehealth triage guideline. The second, RFV Type Intake, starts by asking the user the type of problem (Symptom, Symptom - Pain, Injury, or Exposure) and then a small, standardized set of 3 to 7 questions based on the RFV Type.

Methods: We conducted A/B testing across 11 clinical guidelines (Chest Pain, Leg Pain, Leg Swelling and Edema, Dizziness - Lightheadedness, Rash or Redness - Localized, Pregnancy - Abdominal Pain Greater Than 20 Weeks EGA, Rectal Bleeding, Weakness and Fatigue, Abdominal Pain - Male, Diarrhea, and Knee Pain) using 236 validated clinical scenarios and 472 chatbot conversations. Nine guidelines were tested using the same model, chatbot version, and content set (GPT 5.2, Chatbot POC 11.2, 314 guidelines). Two guidelines (Diarrhea and Knee Pain) were evaluated using a prior model and version. A nurse triage expert used each clinical scenario as a script and completed one conversation per scenario per intake flow. The chatbot applied Triage Assessment Question (TAQ) logic from the Schmitt-Thompson Clinical Content (STCC) adult after-hours telehealth triage guidelines, which served as the clinical reference standard. We measured disposition accuracy (agreement with validated scenario dispositions benchmarked by five expert nurse telehealth triage leaders) and conversational efficiency (question count, excluding demographic questions).

Results: Across all 11 guidelines (236 scenarios, 472 conversations), the IAQ Intake asked approximately 32% more questions on average (21.0 vs. 15.9 questions per conversation, a difference of 5.1 questions) while producing comparable disposition accuracy. Within a nine-guideline subset tested on the same model and version (195 scenarios, 390 conversations), disposition accuracy was identical at 90.8% for both intake methods. Across all 11 guidelines, the RFV Type Intake achieved equal or higher accuracy on eight of eleven guidelines and required fewer questions on every guideline tested. The IAQ Intake demonstrated higher accuracy on three guidelines (Chest Pain: 88% vs. 80%; Pregnancy - Abdominal Pain: 95.5% vs. 86.4%; Weakness and Fatigue: 85% vs. 80%), most notably Chest Pain, where its broader upfront questioning captured clinically discriminating features (radiation, duration, quality, diaphoresis) that drive high-acuity triage rules. For Leg Pain, the RFV Type Intake achieved higher accuracy (100% vs. 90.5%), possibly because narrower questioning avoided surfacing triage rules that were subsequently misapplied to negative findings. Knee Pain showed identical low accuracy (45%) under both approaches, pointing to an LLM agent challenge interpreting the TAQs, rather than intake method.

Conclusions: A structured RFV Type Intake matched the disposition accuracy of a comprehensive IAQ Intake while substantially improving conversational efficiency. These findings suggest that LLM-based triage chatbots may achieve comparable accuracy with fewer questions when deterministic intake processes constrain the problem space before clinical reasoning is applied. The RFV Type approach may also reduce errors caused when the chatbot asks about findings the patient does not have, which can inadvertently activate triage rules that are then incorrectly applied.

Citation

O’Bryan L, Thompson DA. Tancock K. Comparing Two Patient Intake Approaches for a Conversational Triage Chatbot: A Structured Reason For Visit - Type Intake Versus Initial Assessment Questions. September 3, 2026.