Automatic identification of variables in epidemiological datasets using logic regression.
| dc.contributor.author | Lorenz, Matthias W | |
| dc.contributor.author | Abdi, Negin Ashtiani | |
| dc.contributor.author | Scheckenbach, Frank | |
| dc.contributor.author | Pflug, Anja | |
| dc.contributor.author | Bülbül, Alpaslan | |
| dc.contributor.author | Catapano, Alberico L | |
| dc.contributor.author | Agewall, Stefan | |
| dc.contributor.author | Ezhov, Marat | |
| dc.contributor.author | Bots, Michiel L | |
| dc.contributor.author | Kiechl, Stefan | |
| dc.contributor.author | Orth, Andreas | |
| dc.contributor.author | PROG-IMT study group | |
| dc.date.accessioned | 2024-06-11T13:48:19Z | |
| dc.date.available | 2024-06-11T13:48:19Z | |
| dc.date.issued | 2017-04 | |
| dc.description.abstract | BackgroundFor an individual participant data (IPD) meta-analysis, multiple datasets must be transformed in a consistent format, e.g. using uniform variable names. When large numbers of datasets have to be processed, this can be a time-consuming and error-prone task. Automated or semi-automated identification of variables can help to reduce the workload and improve the data quality. For semi-automation high sensitivity in the recognition of matching variables is particularly important, because it allows creating software which for a target variable presents a choice of source variables, from which a user can choose the matching one, with only low risk of having missed a correct source variable.MethodsFor each variable in a set of target variables, a number of simple rules were manually created. With logic regression, an optimal Boolean combination of these rules was searched for every target variable, using a random subset of a large database of epidemiological and clinical cohort data (construction subset). In a second subset of this database (validation subset), this optimal combination rules were validated.ResultsIn the construction sample, 41 target variables were allocated on average with a positive predictive value (PPV) of 34%, and a negative predictive value (NPV) of 95%. In the validation sample, PPV was 33%, whereas NPV remained at 94%. In the construction sample, PPV was 50% or less in 63% of all variables, in the validation sample in 71% of all variables.ConclusionsWe demonstrated that the application of logic regression in a complex data management task in large epidemiological IPD meta-analyses is feasible. However, the performance of the algorithm is poor, which may require backup strategies. | |
| dc.identifier | 10.1186/s12911-017-0429-1 | |
| dc.identifier.issn | 1472-6947 | |
| dc.identifier.issn | 1472-6947 | |
| dc.identifier.uri | ||
| dc.language | eng | |
| dc.publisher | Springer Science and Business Media LLC | |
| dc.relation.ispartof | BMC medical informatics and decision making | |
| dc.relation.isversionof | 10.1186/s12911-017-0429-1 | |
| dc.rights.uri | ||
| dc.subject | PROG-IMT study group | |
| dc.subject | Humans | |
| dc.subject | Carotid Artery Diseases | |
| dc.subject | Prognosis | |
| dc.subject | Logistic Models | |
| dc.subject | Predictive Value of Tests | |
| dc.subject | Epidemiologic Factors | |
| dc.subject | Algorithms | |
| dc.subject | Databases, Factual | |
| dc.subject | Medical Informatics Applications | |
| dc.subject | Meta-Analysis as Topic | |
| dc.subject | Data Mining | |
| dc.subject | Carotid Intima-Media Thickness | |
| dc.title | Automatic identification of variables in epidemiological datasets using logic regression. | |
| dc.type | Journal article | |
| pubs.begin-page | 40 | |
| pubs.issue | 1 | |
| pubs.organisational-group | Duke | |
| pubs.organisational-group | School of Medicine | |
| pubs.organisational-group | Basic Science Departments | |
| pubs.organisational-group | Biostatistics & Bioinformatics | |
| pubs.organisational-group | Biostatistics & Bioinformatics, Division of Biostatistics | |
| pubs.publication-status | Published | |
| pubs.volume | 17 |
Files
Original bundle
- Name:
- Automatic identification of variables in epidemiological datasets using logic regression.pdf
- Size:
- 1.06 MB
- Format:
- Adobe Portable Document Format