A digital health program has collected two years of symptom scores, medication logs and free-text diaries. Thousands of entries sit in the database. Somebody asks whether the company now has real-world evidence.
It has real data from real people. That is not the same answer.
In brief: real-world data are routinely collected data about patient health or health care. Real-world evidence is the clinical evidence produced by analyzing suitable real-world data with a defined question and an appropriate design. Collection is the beginning, not the conclusion.
FDA defines real-world data as data relating to patient health status or the delivery of health care routinely collected from sources such as electronic health records, claims, registries and digital health technologies. Real-world evidence is clinical evidence about the use and potential benefits or risks of a medical product derived from analysis of those data. Read the FDA real-world evidence overview.
When does patient-reported data become real-world evidence?
The same dataset may be useful for one purpose and inadequate for another.
A treatment diary might describe how often app users report taking a medication. It may not establish the medication's effectiveness. A symptom tracker may show change over time among people who keep using the app. It may not show what happened to everyone who started treatment.
| What the database contains | A question it may help answer | A conclusion it cannot support by itself |
|---|---|---|
| Daily self-reported symptom severity | How symptoms changed among participants who submitted entries | That a treatment caused the change |
| Medication marked taken or skipped | How users reported following a planned routine | Verified exposure or pharmacologic effect |
| Quality-of-life questionnaires | Patient-reported function or burden at measured time points | Population-level benefit without understanding selection and missingness |
| Device-use records | Patterns of recorded use and reasons for non-use | Technical performance when connectivity or logging is incomplete |
| Free-text journals | Themes and experiences that may warrant structured study | Incidence, causality or prevalence without a defensible denominator and method |
This is not a reason to dismiss patient-reported data. It is a reason to match claims to what the data can actually support.
Patient-entered does not automatically mean real-world
A person can enter data on a phone in a traditional randomized trial, a registry, a clinical care program, a consumer health app or a post-market study. The interface may look similar while the data context is different.
Data collected inside a conventional clinical trial are not transformed into RWD solely because a participant reports them remotely. Conversely, patient-generated data collected routinely through registries, care programs, devices or digital health technologies may contribute to RWD.
The study team should preserve the context:
- why the data were collected;
- who was eligible to provide them;
- how participants entered or left the data source;
- what instructions they received;
- which measure or field version was used;
- when and through which device the entry occurred;
- whether the value was edited, derived or transformed; and
- what related information was not collected.
Without that provenance, a clean table can still be scientifically ambiguous.
What makes patient-reported data fit for purpose?
Vendors often describe data as "regulatory grade," "research ready" or "RWE-ready." Those labels can obscure the real test: are the data sufficiently relevant and reliable for the specific question and intended use?
FDA's December 2025 guidance for medical devices explains how it evaluates RWD quality for generating RWE in regulatory decision-making and supersedes the earlier 2017 guidance. Read the current FDA device RWE guidance.
For drugs and biologics, FDA has separate final guidance on assessing registries and on general considerations for using RWD and RWE in regulatory decision-making. Review FDA's registry guidance for drugs and biologics.
Useful wording: a platform can support data collection, provenance, quality controls and export. It cannot guarantee that a future analysis will be accepted as evidence for a specific regulatory, reimbursement or clinical decision.
The seven questions to answer before calling it an RWE program
| Question | Why it matters |
|---|---|
| What is the estimand or research question? | Determines the population, exposure, outcome, time period and analysis that matter |
| Who is represented? | Shows whether app users, registry members or program enrollees differ from the intended population |
| How are exposure and outcomes defined? | Separates self-report, prescription, dispensing, device data and clinician-confirmed events |
| What is missing, and why? | Missing entries may be related to health, burden, disengagement, technology or program exit |
| Which biases and confounders are plausible? | Prevents a temporal association from being presented as treatment effect |
| Can data be linked and audited? | Supports reconciliation with clinical, claims, registry, device or study records where permitted |
| Is the analysis prespecified and reproducible? | Distinguishes planned evidence generation from searching a large dataset for an appealing result |
Why do missing patient-reported data matter?
Patient-reported datasets often become less complete over time. The people who continue logging may differ from those who stop. A person may skip an entry because they feel well, feel unwell, are busy, have lost access or no longer see value in the program.
The system should distinguish:
- an assessment that was not scheduled;
- one that was scheduled but never opened;
- one started but not completed;
- a participant who withdrew;
- a technical delivery or synchronization failure;
- a response declined by the participant; and
- a valid zero or normal value.
A dashboard completion percentage is useful operationally. An evidence program also needs to understand the pattern and likely mechanism of missingness.
Patient-reported outcomes need measurement discipline
PRO data can add information that claims and EHRs often miss: symptoms, function, quality of life and treatment burden.
The team still needs to know whether the instrument measures the concept of interest in the target population, whether the recall period fits the use, whether translations and licenses are appropriate, how scores are derived and what constitutes a meaningful change.
FDA's 2025 final patient-focused drug development guidance addresses the selection and development of fit-for-purpose clinical outcome assessments. Read the FDA COA guidance.
A custom weekly question may still be valuable for program operations. It should not be described as a validated outcome measure unless the supporting work exists.
Analysis creates evidence; the app creates observations
The app's job is to collect the observation faithfully and preserve the surrounding metadata. The evidence program then defines the cohort, variables, comparators, time zero, follow-up, censoring, confounders, sensitivity analyses and interpretation.
The March 2026 final ICH M14 guidance covers planning, designing, analyzing and reporting non-interventional studies that use RWD for medicine safety assessment. It emphasizes the research question, data-source selection, variables, bias, confounding, analysis and reporting. Read FDA's final ICH M14 guidance.
Those responsibilities do not disappear when the source data are prospective, digital or patient-generated.
For patients who find this article: information you record about symptoms and treatment experience can help researchers understand health outside clinic visits. A responsible program should explain why the data are collected, whether they are used for care, research or product improvement, and how your privacy and choices are handled.
AI can structure data without solving the study design
AI may help classify free text, normalize terms, identify possible duplicate records or prepare a timeline for human review.
Those transformations need validation, provenance and quality controls for their intended use. AI does not remove selection bias, confounding, missing data or ambiguity about exposure.
A model can make unstructured data easier to analyze. It cannot turn a vague objective into a causal question or make an unrepresentative cohort representative.
How can CareClinic support patient-reported data and real-world evidence?
CareClinic provides a participant-facing foundation for prospective longitudinal collection of symptoms, medications, treatments, measurements, mood, sleep, activity, journals, assessments, care plans and other patient-reported information.
That may support an RWD program when the organization needs recurring patient input outside routine clinic documentation. CareClinic can also be scoped for branded or study-specific participant experiences and structured exports.
CareClinic should not be described as automatically generating regulatory or publication-ready RWE. The research question, cohort, consent, governance, data linkage, variable definitions, quality controls, statistical analysis and interpretation remain part of the evidence program.
Related CareClinic guides
- The Device Launched. The Follow-Up Didn't.
- A Rare Disease Registry Is More Than a Signup Form
- A Reminder Is Not a Patient Support Program
Start with the question you want the data to answer
Tell the CareClinic team which population you need to follow, what patients report, how often data are collected, which other sources need linking and how the resulting dataset will be analyzed or exported.
Discuss a Patient-Reported RWD Program
Choose Patient Reported Outcomes (ePRO) for structured longitudinal data collection. Choose OEM or White Label when the program also needs a branded patient experience. Do not include patient information in the inquiry.
Frequently asked questions
What is the difference between real-world data and real-world evidence?
Real-world data are routinely collected data relating to patient health status or health care delivery. Real-world evidence is clinical evidence about the use, benefits or risks of a medical product derived from analysis of real-world data.
Are patient-reported data real-world data?
It can be, depending on how and why it is collected. Patient-generated data collected routinely through registries, digital health technologies or care programs may be a real-world data source. Data collected inside a traditional clinical trial are not made real-world simply because a patient enters them on a phone.
Does a large patient dataset automatically become evidence?
No. Evidence requires a defined question, relevant and reliable data, an appropriate design and analysis, and a transparent account of missingness, bias, confounding and limitations.
Can ePRO data support real-world evidence?
Potentially. ePRO may contribute patient-centered outcomes to an RWE study when the measure, population, timing, provenance, completeness and analysis are fit for the intended use.
What makes patient-reported data fit for purpose?
Fit-for-purpose data are sufficiently relevant and reliable for the specific research question. Important considerations include who is represented, what was measured, how consistently it was collected, source and timing, missing data, validation, linkage and quality controls.
Can app data support an FDA submission?
Potentially, but acceptance depends on the specific regulatory context, study design, data quality and applicable guidance. An app vendor cannot guarantee that collected data will be accepted for a submission.
Can CareClinic generate real-world evidence?
CareClinic can support prospective longitudinal data collection and patient-reported information. Turning those data into RWE requires a research question, protocol, governance, data-quality plan, analysis and appropriate scientific and regulatory oversight.
Can AI turn unstructured patient reports into RWE?
AI may help structure or classify text, but the transformation must be validated for the intended use and preserve provenance. AI does not remove selection bias, confounding, missing data or the need for a defensible study design.
Should a small biotech start collecting data before it has an RWE protocol?
It can collect operational or exploratory data, but it should not assume those records will later answer a regulatory or scientific question. Defining likely future uses early improves consent, variables, timing, provenance and data quality.
Educational information only. This article is not regulatory, legal, clinical, statistical, epidemiologic, privacy or research advice. Organizations should assess their question, data sources, protocol, analysis and intended use with the appropriate professional advisers.


