ClusterImportance#

class hidimstat.ClusterImportance(vim, clustering)[source]#

Bases: BaseVariableImportance

Clustered inference with any variable importance method.

This algorithm computes a single clustered inference on groups of features using an arbitrary variable importance measure for statistical inference.

Parameters:
vim: hidimstat.BaseVariableImportance

An instance of any variable importance method that derives from hidimstat’s BaseVariableImportance.

clustering: sklearn.cluster.FeatureAgglomeration

An instance of a clustering method that operates on features.

Attributes:
vim_hidimstat.BaseVariableImportance

Fitted variable importance method that will be clustered.

clustering_sklearn.cluster.FeatureAgglomeration

Fitted clustering object.

clustering_samples_ndarray, (n_samples*cluster_frac,)

Indices of samples used for clustering.

importances_ndarray, shape (n_clusters,) or (n_clusters, n_tasks)

Estimated coefficients at cluster level.

pvalues_ndarray, shape (n_clusters,)

P-values for each cluster.

n_features_int

Number of features in the original data.

Notes

Added in version 0.4.0.

__init__(vim, clustering)[source]#
fit(X, y)[source]#

Fit the clustering and variable importance method on the data.

Parameters:
Xndarray, shape (n_samples, n_features)

Input data matrix.

yndarray, shape (n_samples,) or (n_samples, n_tasks)

Target variable(s).

Returns:
selfCluVI

Fitted estimator.

importance(X, y)[source]#

Compute feature importance from the underlying variable importance method. Then map the importance scores from cluster level back to feature level.

Parameters:
Xndarray, shape (n_samples, n_features)

Input data matrix.

yndarray, shape (n_samples,) or (n_samples, n_tasks)

Target variable(s).

fit_importance(X, y)[source]#

Fit the model and compute feature importance.

Parameters:
Xndarray, shape (n_samples, n_features)

Input data matrix.

yndarray, shape (n_samples,) or (n_samples, n_tasks)

Target variable(s).

Returns:
importances_ndarray, shape (n_features,) or (n_features, n_tasks)

Estimated importance values at feature level.

fdr_selection(fdr, fdr_control='bhq', reshaping_function=None, two_tailed_test=False)[source]#

Performs feature selection based on False Discovery Rate (FDR) control.

Parameters:
fdrfloat

The target false discovery rate level (between 0 and 1)

fdr_control: {‘bhq’, ‘bhy’}, default=’bhq’

The FDR control method to use: - ‘bhq’: Benjamini-Hochberg procedure - ‘bhy’: Benjamini-Hochberg-Yekutieli procedure

reshaping_function: callable or None, default=None

Optional reshaping function for FDR control methods. If None, defaults to sum of reciprocals for ‘bhy’.

two_tailed_test: bool, default=False

If True, performs two-tailed test selection using both p-values for positive effects and one-minus p-values for negative effects. The sign of the effect is determined from the sign of the importance scores.

Returns:
selectedndarray of int of shape (n_features,)

Integer array indicating the selected features. 1 indicates selected features with positive effects, -1 indicates selected features with negative effects, 0 indicates non-selected features.

Raises:
ValueError

If importances_ haven’t been computed yet

AssertionError

If pvalues_ are missing or fdr_control is invalid

fwer_selection(fwer, procedure='bonferroni', n_tests=None, two_tailed_test=False)[source]#

Performs feature selection based on Family-Wise Error Rate (FWER) control.

Parameters:
fwerfloat

The target family-wise error rate level (between 0 and 1)

procedure{‘bonferroni’}, default=’bonferroni’

The FWER control method to use: - ‘bonferroni’: Bonferroni correction

n_testsint or None, default=None

Factor for multiple testing correction. If None, uses the number of clusters or the number of features in this order.

two_tailed_testbool, default=False

If True, uses the sign of the importance scores to indicate whether the selected features have positive or negative effects.

Returns:
selectedndarray of int of shape (n_features,)

Integer array indicating the selected features. 1 indicates selected features with positive effects, -1 indicates selected features with negative effects, 0 indicates non-selected features.

get_metadata_routing()[source]#

Get metadata routing of this object.

Please check User Guide on how the routing mechanism works.

Returns:
routingMetadataRequest

A MetadataRequest encapsulating routing information.

get_params(deep=True)[source]#

Get parameters for this estimator.

Parameters:
deepbool, default=True

If True, will return the parameters for this estimator and contained subobjects that are estimators.

Returns:
paramsdict

Parameter names mapped to their values.

importance_selection(k_best=None, percentile=None, threshold_max=None, threshold_min=None)[source]#

Selects features based on variable importance.

Parameters:
k_bestint, default=None

Selects the top k features based on importance scores.

percentilefloat, default=None

Selects features based on a specified percentile of importance scores.

threshold_maxfloat, default=None

Selects features with importance scores below the specified maximum threshold.

threshold_minfloat, default=None

Selects features with importance scores above the specified minimum threshold.

Returns:
selectionarray-like of shape (n_features,)

Binary array indicating the selected features.

plot_importance(ax=None, ascending=False, feature_names=None, **seaborn_barplot_kwargs)[source]#

Plot feature importances as a horizontal bar plot.

Parameters:
axmatplotlib.axes.Axes or None, (default=None)

Axes object to draw the plot onto, otherwise uses the current Axes.

ascending: bool, default=False

Whether to sort features by ascending importance.

**seaborn_barplot_kwargsadditional keyword arguments

Additional arguments passed to seaborn.barplot. https://seaborn.pydata.org/generated/seaborn.barplot.html

Returns:
axmatplotlib.axes.Axes

The Axes object with the plot.

pvalue_selection(k_lowest=None, percentile=None, threshold_max=0.05, threshold_min=None, alternative_hypothesis=False)[source]#

Selects features based on p-values.

Parameters:
k_lowestint, default=None

Selects the k features with lowest p-values.

percentilefloat, default=None

Selects features based on a specified percentile of p-values.

threshold_maxfloat, default=0.05

Selects features with p-values below the specified maximum threshold (0 to 1).

threshold_minfloat, default=None

Selects features with p-values above the specified minimum threshold (0 to 1).

alternative_hypothesisbool, default=False

If True, selects based on 1-pvalues instead of p-values.

Returns:
selectionarray-like of shape (n_features,)

Binary array indicating the selected features (True for selected).

set_params(**params)[source]#

Set the parameters of this estimator.

The method works on simple estimators as well as on nested objects (such as Pipeline). The latter have parameters of the form <component>__<parameter> so that it’s possible to update each component of a nested object.

Parameters:
**paramsdict

Estimator parameters.

Returns:
selfestimator instance

Estimator instance.

Examples using hidimstat.ClusterImportance#

Support Recovery on fMRI Data

Support Recovery on fMRI Data

Ensemble Clustered Inference on 2D Data

Ensemble Clustered Inference on 2D Data