Preserving medical correctness, readability and consistency in de-identified health records
- Kostas Pantazos,
- ,
- Søren Lippert
Research Output:
Journal Article or Conference Article in Journal
Journal article
Peer-reviewOpen access
Publication Information
Output type
Research Output:
Journal Article or Conference Article in Journal
Journal article
Peer-reviewOriginal language
EnglishPages from-to (Number of pages)
Pages 1-13 (13 pages)Journal (Volume, Issue Number)
Health Informatics JournalPublication milestones
- Published - 01/01/2016
Publication status
Published - 01/01/2016
ISSN
1460-4582Publication IDs
- Scopus: 85034637745
Abstract
A health record database contains structured data fields that identify the patient, such as patient ID, patient
name, e-mail and phone number. These data are fairly easy to de-identify, that is, replace with other
identifiers. However, these data also occur in fields with doctors’ free-text notes written in an abbreviated
style that cannot be analyzed grammatically. If we replace a word that looks like a name, but isn’t, we degrade
readability and medical correctness. If we fail to replace it when we should, we degrade confidentiality. We de-identified an existing Danish electronic health record database, ending up with 323,122 patient health records. We had to invent many methods for de-identifying potential identifiers in the free-text notes. The de-identified health records should be used with caution for statistical purposes because we removed health records that were so special that they couldn’t be de-identified. Furthermore, we distorted geography by replacing zip codes with random zip codes.
name, e-mail and phone number. These data are fairly easy to de-identify, that is, replace with other
identifiers. However, these data also occur in fields with doctors’ free-text notes written in an abbreviated
style that cannot be analyzed grammatically. If we replace a word that looks like a name, but isn’t, we degrade
readability and medical correctness. If we fail to replace it when we should, we degrade confidentiality. We de-identified an existing Danish electronic health record database, ending up with 323,122 patient health records. We had to invent many methods for de-identifying potential identifiers in the free-text notes. The de-identified health records should be used with caution for statistical purposes because we removed health records that were so special that they couldn’t be de-identified. Furthermore, we distorted geography by replacing zip codes with random zip codes.
Publication metrics
PlumX, opens in new tab
Captures
31
Citations
16
Access to documents
Accepted author manuscript, 371.63 KB
