Automatic identification of variables in epidemiological datasets using logic regression.

dc.contributor.author

Lorenz, Matthias W

dc.contributor.author

Abdi, Negin Ashtiani

dc.contributor.author

Scheckenbach, Frank

dc.contributor.author

Pflug, Anja

dc.contributor.author

Bülbül, Alpaslan

dc.contributor.author

Catapano, Alberico L

dc.contributor.author

Agewall, Stefan

dc.contributor.author

Ezhov, Marat

dc.contributor.author

Bots, Michiel L

dc.contributor.author

Kiechl, Stefan

dc.contributor.author

Orth, Andreas

dc.contributor.author

PROG-IMT study group

dc.date.accessioned

2024-06-11T13:48:19Z

dc.date.available

2024-06-11T13:48:19Z

dc.date.issued

2017-04

dc.description.abstract

Background

For an individual participant data (IPD) meta-analysis, multiple datasets must be transformed in a consistent format, e.g. using uniform variable names. When large numbers of datasets have to be processed, this can be a time-consuming and error-prone task. Automated or semi-automated identification of variables can help to reduce the workload and improve the data quality. For semi-automation high sensitivity in the recognition of matching variables is particularly important, because it allows creating software which for a target variable presents a choice of source variables, from which a user can choose the matching one, with only low risk of having missed a correct source variable.

Methods

For each variable in a set of target variables, a number of simple rules were manually created. With logic regression, an optimal Boolean combination of these rules was searched for every target variable, using a random subset of a large database of epidemiological and clinical cohort data (construction subset). In a second subset of this database (validation subset), this optimal combination rules were validated.

Results

In the construction sample, 41 target variables were allocated on average with a positive predictive value (PPV) of 34%, and a negative predictive value (NPV) of 95%. In the validation sample, PPV was 33%, whereas NPV remained at 94%. In the construction sample, PPV was 50% or less in 63% of all variables, in the validation sample in 71% of all variables.

Conclusions

We demonstrated that the application of logic regression in a complex data management task in large epidemiological IPD meta-analyses is feasible. However, the performance of the algorithm is poor, which may require backup strategies.
dc.identifier

10.1186/s12911-017-0429-1

dc.identifier.issn

1472-6947

dc.identifier.issn

1472-6947

dc.identifier.uri

https://hdl.handle.net/10161/31168

dc.language

eng

dc.publisher

Springer Science and Business Media LLC

dc.relation.ispartof

BMC medical informatics and decision making

dc.relation.isversionof

10.1186/s12911-017-0429-1

dc.rights.uri

https://creativecommons.org/licenses/by-nc/4.0

dc.subject

PROG-IMT study group

dc.subject

Humans

dc.subject

Carotid Artery Diseases

dc.subject

Prognosis

dc.subject

Logistic Models

dc.subject

Predictive Value of Tests

dc.subject

Epidemiologic Factors

dc.subject

Algorithms

dc.subject

Databases, Factual

dc.subject

Medical Informatics Applications

dc.subject

Meta-Analysis as Topic

dc.subject

Data Mining

dc.subject

Carotid Intima-Media Thickness

dc.title

Automatic identification of variables in epidemiological datasets using logic regression.

dc.type

Journal article

pubs.begin-page

40

pubs.issue

1

pubs.organisational-group

Duke

pubs.organisational-group

School of Medicine

pubs.organisational-group

Basic Science Departments

pubs.organisational-group

Biostatistics & Bioinformatics

pubs.organisational-group

Biostatistics & Bioinformatics, Division of Biostatistics

pubs.publication-status

Published

pubs.volume

17

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Automatic identification of variables in epidemiological datasets using logic regression.pdf
Size:
1.06 MB
Format:
Adobe Portable Document Format