FAQ chatbots for patients: what they answer, where they stop
Which questions a hospital FAQ chatbot can take, where it hands over to staff, how answers are grounded in approved documents, and what to check before buying.
A frequently-asked-questions (FAQ) chatbot can take part of the questions patients ask every day at the front desk and on the phone: opening hours, documents to bring, how to prepare for a test, prices, how to get to the hospital. Its value depends on two decisions made before launch: what it is allowed to answer, and where its answers come from. This article sets out what published data say about those two decisions and what a hospital should check before buying.
How much of what patients ask is administrative
A 2026 study posted on medRxiv analysed 30,390 patient portal messages from 4,817 patients of a United States ophthalmology clinic, sent between June 2014 and July 2024 (Kim et al., 2026). According to the authors, nearly half of all messages addressed administrative issues: scheduling, medication refills and insurance. The rest were clinical, and among those the leading topics were vision disturbances (20.8%), glaucoma-related symptoms (8.7%), imaging or tumour questions (7.5%) and postoperative concerns (7.4%).
A study published in JAMA in June 2026, using the Epic Cosmos network (more than 2,000 hospitals and 47,000 clinics), found that patient-authored messages rose 153% between 2020 and 2025, from 0.99 to 2.50 messages per active patient per year (Long et al., 2026). Over the study period there were 1.34 billion patient-authored messages and 1.59 billion telephone encounters. The written channel did not replace the telephone; it was added alongside it.
What it can answer safely
A well-bounded FAQ chatbot answers questions whose answer is the same for every patient and has been approved by the institution:
- opening hours of departments, the laboratory and the cashier;
- documents required for admission or a consultation;
- preparation for tests and investigations, using the exact wording of the internal protocol;
- the price list for paid services;
- how to reach the hospital, where to park, which entrance serves which building;
- insurance questions: which patient categories are covered, which documents are required.
The European guidance MDCG 2019-11 on qualifying software under the Medical Device Regulation (MDR) states that software which only stores, archives, communicates or performs a simple search of information is not a medical device (European Commission, MDCG 2019-11). A chatbot that repeats the opening hours and the list of documents falls into that category. What matters is the manufacturer's intended purpose: once the software is meant to interpret data for an individual patient for diagnosis, monitoring or treatment, the discussion changes.
Where it must stop
Three areas stay closed to an FAQ chatbot: symptoms ("I have chest pain, what could it be?"), treatment ("can I take ibuprofen with my anticoagulant?") and medication ("how much of this drug can I take?").
A 2024 experimental study put 10 emergency-care questions to four public chatbots (ChatGPT 3.5, Google Bard, Bing AI Chat and Claude) and had the answers graded by five emergency physicians (Yau et al., 2024). Clarity was good (85%), but accuracy and completeness scored 50%, source reliability 10%, and between 5% and 35% of responses contained dangerous information. The authors conclude that advice on when to seek emergency care was frequently incomplete and inaccurate.
A meta-analysis of 17 studies found an integrated accuracy of 56% (95% confidence interval: 51-60%) for ChatGPT answers to medical questions, with high heterogeneity between studies (Wei et al., 2024). One correct answer in two is not an acceptable standard for clinical advice. The standard reply to any question about symptoms, treatment or medication must therefore always be the same: I cannot answer medical questions; I can book you an appointment or connect you with staff; in an emergency call 112.
| Type of question | FAQ chatbot | Why |
|---|---|---|
| Hours, documents, directions, prices | Answers, from the approved document | Fixed information, no individual interpretation (MDCG 2019-11) |
| Preparation for tests | Answers, with the internal protocol text | Identical for everyone; medication changes stay with the doctor |
| Insurance, contracts, paperwork | Answers | Administrative information |
| Symptoms ("what could it be?") | Stops, hands over to staff | Accuracy 50%, dangerous information in 5-35% of answers (Yau et al., 2024) |
| Treatment, doses, interactions | Stops, hands over to staff | Interpretation for an individual patient = medical purpose (MDCG 2019-11) |
| Complaints, sensitive situations | Stops, hands over to staff | Need for empathy and human accountability |
How answers are grounded: retrieval, not free generation
A large language model (LLM) left on its own answers from what it learned in training, that is, from the public internet. For a hospital, the correct answer to "how much does an abdominal ultrasound cost?" is not on the internet; it is in the price list approved by management. The technical answer is called retrieval-augmented generation (RAG): for each question, the system first searches the institution's documents for the relevant passages, then asks the model to formulate the answer only from those passages, with a reference to the source.
The difference is measurable. In the Almanac study, a RAG system for clinical guidelines was compared with models without retrieval on 130 clinical scenarios rated by five physicians; factuality rose by a mean of 18%, with improvements in completeness and safety as well (Zakka et al., 2023). The study concerns clinical guidelines rather than administrative FAQs, but the mechanism is the same.
The practical rule: if the answer does not exist in an approved hospital document, the chatbot does not invent it. It says it does not know and hands over to a person.
For RAG to work, the institution needs a clean document set, with an owner and a review date for each document. When the laboratory's hours change, the document changes; nothing is "retrained".
Hallucination risk and how to measure it
A hallucination is an answer that sounds convincing but is not supported by the source. Even with RAG, the model can add details that are not in the retrieved passage. How often depends on the model. The Vectara leaderboard uses a dedicated evaluation model (HHEM) to measure how often an LLM introduces false information when summarising a document it has been given, using only the information in that document (Vectara, 2026). At the update of 22 September 2026, rates ranged from 1.8% for the models with the lowest rate to 24.2% for those with the highest.
The leaderboard is a summarisation test, not a test on patient questions, so the figures do not transfer directly; the method does. A hospital can do the same before launch and after every significant change:
- Collect 100-200 real questions, as patients write them, from the telephone log and from e-mail.
- For each, write the expected answer and the source document.
- Run the questions through the chatbot and have staff mark each answer: correct and supported by the source, correct but unsupported, wrong, or correctly handed over.
- Publish internally the rate of wrong answers and the rate of handovers. A wrong answer to an administrative question is a quality problem; an answer to a medical question, even a correct one, is a breach of the boundary.
Handover to staff, 24/7 coverage and languages
Handover is not a secondary feature; it is half of the product. The chatbot must hand over to a person in four situations: a medical question, a question with no source in the documents, a patient who explicitly asks for a human, and any sign of urgency or distress. Handover means a clear message about who will reply and when, transfer of the conversation to a channel with a real person and a record of the case.
Round-the-clock coverage is the real advantage of an FAQ chatbot: answers to administrative questions are available on Sunday evening too, when the patient is preparing documents for Monday's admission. Medical questions received at night get the redirection message and the 112 number.
Languages matter especially in the Republic of Moldova. In the 2024 census, 45% of the population declared Moldovan as the language they usually speak, 33.7% Romanian and 15.9% Russian; Russian is the mother tongue of 11.6% (National Bureau of Statistics, 2025). A chatbot that answers only in Romanian leaves at least one patient in six outside the service. The reasonable set is Romanian, Russian and English, with the same source documents translated and approved in each language, not translated on the fly by the model.
What patients say about chatbots
Acceptance data are mixed. A British mixed-methods study (29 interviews and an online survey of 216 respondents) found moderate acceptability, 67%, for AI-led health chatbots; acceptance was lower among people with poorer perceived IT skills and higher with perceived usefulness and trust (Nadarzynski et al., 2019). The concerns raised were accuracy, cyber-security and the inability to empathise.
In the United States, a Pew survey of 11,004 adults found that 60% would be uncomfortable if their health care provider relied on AI for diagnosis and treatment, 79% would not want to use a chatbot for mental-health support, 75% worry that providers will adopt AI too quickly, and 37% fear for the security of their health records (Pew Research Center, 2023).
Read together, the figures say the same thing as the boundary above: patients accept AI for useful, verifiable tasks and reject it where it touches diagnosis and treatment.
GDPR basics for chat transcripts
Even if the chatbot answers only administrative questions, patients write what they want. A message such as "I have diabetes, do I need to fast for test X?" contains health data. The General Data Protection Regulation (GDPR) treats such data under Article 9 as a special category whose processing is prohibited unless an exception applies, for example explicit consent or necessity for the provision of health care on the basis of law or a contract with a health professional (GDPR, Article 9). Article 5 requires that data be collected for specified purposes, limited to what is necessary and kept in a form that identifies the person for no longer than needed. Article 13 requires that the person be informed about the controller, the purposes, the legal basis and the storage period.
For chat transcripts this translates into a few concrete rules: no account and no identifying data for general questions; a short, documented retention period; automatic masking of telephone numbers, personal identification numbers and addresses in the log; a data processing agreement (DPA) with the vendor that fixes the storage location and prohibits use of transcripts for model training; log access only for designated staff.
The Artificial Intelligence Act (AI Act) adds one more obligation. Article 50 requires that people interacting directly with an AI system be informed of that fact, unless it is obvious from the circumstances; the obligation applies from 2 August 2026 (AI Act, Article 50). The chatbot's first sentence must say that it is an automated assistant.
What this means for a hospital in Romania or Moldova
- Measure your own questions. Two weeks of noting, at the front desk and the switchboard, the reason for each call. If the administrative share is close to half, as in the ophthalmology study, you have a use case.
- Write the source documents. Hours, documents, preparations, prices, directions, insurance, each with an owner and a review date. Without these documents, any chatbot answers from the internet.
- Fix the boundary in writing. The list of forbidden questions and the exact text of the redirection message, approved by the medical director.
- Build the test set of 100-200 real questions and ask the vendor to run it before signing.
- Define the handover. Who receives handed-over conversations, within what time they reply, how the case is closed.
Questions for the vendor, with the expected answer:
- Do answers come exclusively from our documents, with a reference to the source? (Yes, through retrieval; with no retrieved passage there is no answer.)
- Where are transcripts stored and in which country? (In the European Union, under a signed DPA.)
- Are transcripts used to train models? (No.)
- What wrong-answer rate did you measure on our test set, and how is it measured after launch? (A number, a method, a monthly report.)
- How does handover to staff work and what does staff see? (The transcript and the source document used.)
- Which languages does it answer in and who approves the translated texts? (RO/RU/EN, with approved source documents in each language.)
Consdinamic builds software and AI to order, with its deepest specialisation in healthcare, and its portfolio includes an FAQ chatbot, a voice assistant for appointments and QR-code wayfinding, all available in Romanian, Russian and English. Its own products run daily in a private medical network in Romania (Gral Medical, 29 locations), where the review platform has been in use for more than 15 months and took the review answer rate from about 30% to 100%. The approach described here, approved documents as the only source, a fixed medical boundary and handover to staff, is the one we apply.
Conclusion
An FAQ chatbot for patients is not a digital doctor and must not be presented as one. It is a tool that repeats, at any hour and in the patient's language, information the institution has already approved: hours, documents, preparations, prices, directions, insurance. Published data show that nearly half of patient messages fall into this category and that answers anchored in documents are measurably more accurate than freely generated ones. On symptoms, treatment and medication, public chatbots are wrong too often to be allowed to answer. The boundary is set in writing, tested on real questions and checked monthly. The rest is handover to people.
- Kim JY et al., Unseen Insights: An AI-Powered Exploration of Secure Patient Messages in Ophthalmology, medRxiv, 2026 — 30,390 messages from 4,817 patients; nearly half administrative; distribution of clinical topics
- Long JJ et al., Trends in Patient Portal Messages, Office Visits, and Telephone Encounters, JAMA, 2026 — patient-authored messages +153% in 2020-2025, 0.99 to 2.50 per active patient per year; telephone vs message volumes
- Yau JY et al., Accuracy of Prospective Assessments of 4 LLM Chatbot Responses to Patient Questions About Emergency Care, JMIR, 2024 — clarity 85%, accuracy and completeness 50%, sources 10%, dangerous information in 5-35% of responses
- Wei Q et al., Evaluation of ChatGPT-Generated Medical Responses: A Systematic Review and Meta-Analysis, J Biomed Inform, 2024 — integrated accuracy 56% (95% CI: 51-60%) across 17 studies
- Zakka C et al., Almanac: Retrieval-Augmented Language Models for Clinical Medicine, 2023 — 130 clinical scenarios, 5 physician raters, factuality +18% with document retrieval
- Vectara, Hallucination Leaderboard (HHEM), updated 22 September 2026 — hallucination rate when summarising a document: 1.8% to 24.2% depending on the model
- Nadarzynski T et al., Acceptability of AI-led chatbot services in healthcare: a mixed-methods study, Digital Health, 2019 — 29 interviews, 216 survey respondents, 67% acceptability; concerns: accuracy, security, lack of empathy
- Pew Research Center, 60% of Americans Would Be Uncomfortable With Provider Relying on AI in Their Own Health Care, 2023 — 11,004 adults; 60% uncomfortable with AI in diagnosis; 79% would not use a mental-health chatbot; 75% fear adoption is too fast
- National Bureau of Statistics of the Republic of Moldova, Final results of the 2024 Census: ethnocultural characteristics, 2025 — usual spoken language: Moldovan 45%, Romanian 33.7%, Russian 15.9%; Russian mother tongue 11.6%
- Regulation (EU) 2016/679 (GDPR), Articles 5, 9 and 13 — health data is a special category; minimisation, storage limitation, information to the data subject
- Regulation (EU) 2024/1689 (AI Act), Article 50 - transparency obligations — people must be told they are interacting with an AI system; applies from 2 August 2026
- MDCG 2019-11, Guidance on qualification and classification of software (MDR/IVDR), European Commission — software that only stores, communicates or performs simple search is not a medical device; intended purpose decides
Tell us what you need. We come back with a prototype, not with slides.