|
A model for the classification
of 70 HIV-1 protease crystal structure
binding pockets to one of its complexed
FDA approved protease inhibitors
utilizing a hybrid data mining modeling
approach has been developed. 456
chemical descriptors have been derived
from the binding pocket structure
of each crystal structure. Two hybrid
approaches were developed, Random
Forest-Linear Discriminant Analysis
(RF-LDA) and Random Forest-Logistic
Regression (RF-LR). Random Forest
is used as a feature selection proxy,
selecting for the most relevant
descriptors used to train its classification
model. The top ranked descriptors
are then used to train the subsequent
LDA and LR models. The classification
performance of LDA and LR are compared
against the Random Forest classifier
used to perform the feature selection.
As a model validation step, hierarchical
clustering of the top ranked descriptors
is performed to verify the descriptor
selection by Random Forest can group
together the binding pocket structures
based on their complexed ligands.
Analysis of the top ranked chemical
descriptors would play a crucial
role in understanding the HIV-1
protease binding pocket in terms
of its drug resistance and protein-ligand
interactions.
|