Classifiers Module#
The ex_fuzzy.classifiers module provides the main classification interface for the ex-fuzzy library.
Overview#
This module contains high-level classifiers built on fuzzy rule mining.
FuzzyRulesClassifier mines fuzzy association rules on a capped feature
space and selects a compact rule base with a genetic algorithm (see
Fuzzy association rule classifier). RuleMineClassifier and
RuleFineTuneClassifier pass mined candidate rules to
BaseFuzzyRulesClassifier.
Classes#
FuzzyRulesClassifier#
- class ex_fuzzy.classifiers.FuzzyRulesClassifier(nRules=None, nAnts=3, fuzzy_type=FUZZY_SETS.t1, tolerance=0.0, verbose=False, n_class=None, runner=1, expansion_factor=1, linguistic_variables=None, max_features=8, feature_selector='mutual_info', feature_selection='per_class', n_linguistic_variables='auto', min_support=0.05, min_confidence=0.5, candidates_per_class=50, rule_mode='additive', n_gen=100, pop_size=60, patience=20, rule_penalty=0.001, rules_per_class=None, random_state=None)[source]#
Bases:
ClassifierMixin,BaseEstimatorFuzzy association rule classifier with feature capping and genetic rule selection.
The classifier follows the FARC-HD family of fuzzy association rule classifiers (Alcalá-Fdez, Alcalá and Herrera, IEEE Transactions on Fuzzy Systems 19(5), 2011), which learns a compact linguistic rule base:
The input space is capped to the
max_featuresmost informative features, which keeps rule generation tractable and removes noise.Each kept feature gets a fixed fuzzy partition (quantile-based terms for numerical features, one term per category for categorical ones).
Candidate rules with up to
nAntsconditions are mined per class and filtered by support, confidence and penalized certainty factor.A covering-based subgroup-discovery prescreen keeps a diverse pool of up to
candidates_per_classrules per class.A genetic algorithm selects a compact subset of rules that maximizes training accuracy with a small penalty per rule, capped by
nRulesorrules_per_classwhen set.
Rules are weighted by their penalized certainty factor and combined as in Ex-Fuzzy’s regression: with
rule_mode="additive"every matching rule votes with its weighted firing, and withrule_mode="sufficient"only each sample’s strongest weighted rule decides. Samples that fire no rule take the training majority class.ex_fuzzy.BaseFuzzyRulesClassifieris not modified.- Parameters:
nRules (int or None, default=None) – Maximum number of rules in the final rule base.
Noneremoves the absolute cap.nAnts (int or "auto", default=3) – Maximum number of conditions per rule.
"auto"allows four conditions when a class’s rules are mined from at most six features, where interactions must carry the model, and three otherwise. Four or five conditions are supported; the candidate search is then pruned bymin_support, so raising it also speeds up long rules.fuzzy_type (FUZZY_SETS, default=FUZZY_SETS.t1) – Only Type-1 fuzzy sets are supported.
tolerance (float, default=0.0) – Minimum penalized certainty factor a candidate rule must exceed.
verbose (bool, default=False) – Print a summary of the fitting stages.
n_class (int, optional) – Ignored; the classes are read from
y. Kept for compatibility.runner (int, default=1) – Ignored; the rule selection is vectorized. Kept for compatibility.
expansion_factor (int, default=1) – Multiplies
candidates_per_class, enlarging the pool the genetic search chooses from.linguistic_variables (list of fuzzyVariable, optional) – One fuzzy variable per input feature. When omitted, partitions are built from the data.
max_features (int, "auto" or None, default=8) – Maximum number of input features to keep.
Nonekeeps them all."auto"keeps 8 for problems with up to five classes and 16 for problems with more, which need more features to separate.feature_selector ({"mutual_info", "f_classif"} or callable, default="mutual_info") – Feature relevance score. A callable receives
(X, y)and returns one score per feature.feature_selection ({"global", "per_class"}, default="per_class") –
"global"keeps themax_featuresmost relevant features for all classes."per_class"scores features one class against the rest and mines each class’s rules on its ownmax_featuresfeatures, so the rule base can use more features in total while each rule search stays small.n_linguistic_variables (int or "auto", default="auto") – Number of fuzzy terms per numerical feature.
"auto"uses three terms for binary problems and five otherwise.min_support (float, default=0.05) – Minimum class support of a candidate rule.
min_confidence (float, default=0.5) – Minimum confidence of a candidate rule.
candidates_per_class (int, default=50) – Maximum rules per class kept by the prescreen.
rule_mode ({"additive", "sufficient"}, default="additive") – How rules are combined into a class decision.
"additive"sums each class’s weighted rule firing;"sufficient"keeps only each sample’s strongest weighted rule, as inex_fuzzy.BaseFuzzyRulesRegressor.n_gen (int, default=100) – Maximum generations of the rule-selection search.
pop_size (int, default=60) – Population size of the rule-selection search.
patience (int, default=20) – Generations without improvement before the search stops.
rule_penalty (float, default=1e-3) – Accuracy traded for each selected rule per class.
rules_per_class (int, optional) – Cap the rule base at
rules_per_classtimes the number of classes. WhennRulesis also set, the smaller cap applies.random_state (int, optional) – Seed for feature scoring and the genetic search.
- classes_#
Class labels.
- Type:
np.ndarray
- selected_features_#
Indices of the input features the rules use.
- Type:
np.ndarray
- class_features_#
Input feature indices each class’s rules were mined from.
- Type:
list of np.ndarray
- max_features_#
Resolved feature cap.
- Type:
int or None
- n_conditions_#
Resolved maximum number of conditions per class.
- Type:
list of int
- linguistic_variables_#
Partitions of the selected features.
- Type:
list of fuzzyVariable
- rule_base_#
Selected rules, one rule base per class, weighted by certainty factor.
- Type:
MasterRuleBase or None
- n_rules_#
Number of selected rules.
- Type:
int
- majority_class_#
Index of the class predicted when no rule fires.
- Type:
int
- __init__(nRules=None, nAnts=3, fuzzy_type=FUZZY_SETS.t1, tolerance=0.0, verbose=False, n_class=None, runner=1, expansion_factor=1, linguistic_variables=None, max_features=8, feature_selector='mutual_info', feature_selection='per_class', n_linguistic_variables='auto', min_support=0.05, min_confidence=0.5, candidates_per_class=50, rule_mode='additive', n_gen=100, pop_size=60, patience=20, rule_penalty=0.001, rules_per_class=None, random_state=None)[source]#
- fit(X, y, n_gen=None, pop_size=None, checkpoints=0, **kwargs)[source]#
Mine, prescreen and select the rule base.
- Parameters:
X (array-like of shape (n_samples, n_features)) – Training features.
y (array-like of shape (n_samples,)) – Class labels.
n_gen (int) – int, optional: Override the constructor’s search budget for this fit.
pop_size (int) – int, optional: Override the constructor’s search budget for this fit.
checkpoints (int, default=0) – Ignored; kept for compatibility.
**kwargs –
random_stateandpatienceoverride the constructor values.
- Returns:
- FuzzyRulesClassifier
The fitted estimator.
- predict_proba(X)[source]#
Class probabilities from normalized rule scores.
Samples that fire no rule receive the training class distribution.
- internal_classifier()[source]#
The selected rule base wrapped as a
BaseFuzzyRulesClassifier.The wrapped model expects only the selected features (
X[:, selected_features_]) and uses winning-rule inference without the majority-class fallback; use this estimator’spredictfor the classifier’s own decisions.
- set_fit_request(*, checkpoints='$UNCHANGED$', n_gen='$UNCHANGED$', pop_size='$UNCHANGED$')#
Configure whether metadata should be requested to be passed to the
fitmethod.Note that this method is only relevant when this estimator is used as a sub-estimator within a meta-estimator and metadata routing is enabled with
enable_metadata_routing=True(seesklearn.set_config()). Please check the User Guide on how the routing mechanism works.The options for each parameter are:
True: metadata is requested, and passed tofitif provided. The request is ignored if metadata is not provided.False: metadata is not requested and the meta-estimator will not pass it tofit.None: metadata is not requested, and the meta-estimator will raise an error if the user provides it.str: metadata should be passed to the meta-estimator with this given alias instead of the original name.
The default (
sklearn.utils.metadata_routing.UNCHANGED) retains the existing request. This allows you to change the request for some parameters and not others.Added in version 1.3.
- Parameters:
checkpoints (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for
checkpointsparameter infit.n_gen (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for
n_genparameter infit.pop_size (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for
pop_sizeparameter infit.
- Returns:
self – The updated object.
- Return type:
object
- set_score_request(*, sample_weight='$UNCHANGED$')#
Configure whether metadata should be requested to be passed to the
scoremethod.Note that this method is only relevant when this estimator is used as a sub-estimator within a meta-estimator and metadata routing is enabled with
enable_metadata_routing=True(seesklearn.set_config()). Please check the User Guide on how the routing mechanism works.The options for each parameter are:
True: metadata is requested, and passed toscoreif provided. The request is ignored if metadata is not provided.False: metadata is not requested and the meta-estimator will not pass it toscore.None: metadata is not requested, and the meta-estimator will raise an error if the user provides it.str: metadata should be passed to the meta-estimator with this given alias instead of the original name.
The default (
sklearn.utils.metadata_routing.UNCHANGED) retains the existing request. This allows you to change the request for some parameters and not others.Added in version 1.3.
- Parameters:
sample_weight (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for
sample_weightparameter inscore.- Returns:
self – The updated object.
- Return type:
object
RuleMineClassifier#
- class ex_fuzzy.classifiers.RuleMineClassifier(nRules=30, nAnts=4, fuzzy_type=FUZZY_SETS.t1, tolerance=0.0, verbose=False, n_class=None, runner=1, linguistic_variables=None, n_gen=30, pop_size=50, patience=10, random_state=33)[source]#
Bases:
ClassifierMixin,BaseEstimatorA classifier that works by mining a set of candidate rules with a minimum support, confidence and lift, and then using a genetic algorithm that chooses the optimal combination of those rules.
The main classifier that mines candidate rules and then optimizes them using genetic algorithms.
- __init__(nRules=30, nAnts=4, fuzzy_type=FUZZY_SETS.t1, tolerance=0.0, verbose=False, n_class=None, runner=1, linguistic_variables=None, n_gen=30, pop_size=50, patience=10, random_state=33)[source]#
Inits the optimizer with the corresponding parameters.
- Parameters:
nRules (int) – number of rules to optimize.
nAnts (int) – max number of antecedents to use.
fuzzy_type (FUZZY_SETS) – FUZZY_SET enum type in fuzzy_sets module. The kind of fuzzy set used.
tolerance (float) – tolerance for the support/dominance score of the rules.
verbose – if True, prints the progress of the optimization.
n_class (int) – number of classes in the problem. If None (default) the classifier will compute it empirically.
runner (int) – number of threads to use.
linguistic_variables (list[fuzzyVariable]) – linguistic variables per antecedent.
n_gen (int) – number of generations of the genetic search. fit can override it for one call.
pop_size (int) – population size of the genetic search. fit can override it for one call.
patience (int | None) – generations without improvement before the search stops early; None runs every generation.
random_state (int) – random seed of the genetic search.
- property fl_classifier: BaseFuzzyRulesClassifier#
The classifier that performs the final predictions. Built when fit is called.
- fit(X, y, n_gen=CONSTRUCTOR, pop_size=CONSTRUCTOR, **kwargs)[source]#
Trains the model with the given data.
- Parameters:
X (array) – samples to train.
y (array) – labels for each sample.
n_gen (int) – number of generations to compute in the genetic optimization. Defaults to the constructor’s value.
pop_size (int) – number of subjects per generation. Defaults to the constructor’s value.
kwargs – additional parameters for the genetic optimization, including early stopping with patience=10 and min_delta=1e-4 by default. See fit method in BaseRuleBaseClassifier.
- Returns:
the fitted classifier.
- predict(X)[source]#
Predict for each sample the corresponding class.
- Parameters:
X (array) – samples to predict.
- Returns:
a class for each sample.
- Return type:
array
- set_fit_request(*, n_gen='$UNCHANGED$', pop_size='$UNCHANGED$')#
Configure whether metadata should be requested to be passed to the
fitmethod.Note that this method is only relevant when this estimator is used as a sub-estimator within a meta-estimator and metadata routing is enabled with
enable_metadata_routing=True(seesklearn.set_config()). Please check the User Guide on how the routing mechanism works.The options for each parameter are:
True: metadata is requested, and passed tofitif provided. The request is ignored if metadata is not provided.False: metadata is not requested and the meta-estimator will not pass it tofit.None: metadata is not requested, and the meta-estimator will raise an error if the user provides it.str: metadata should be passed to the meta-estimator with this given alias instead of the original name.
The default (
sklearn.utils.metadata_routing.UNCHANGED) retains the existing request. This allows you to change the request for some parameters and not others.Added in version 1.3.
- Parameters:
n_gen (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for
n_genparameter infit.pop_size (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for
pop_sizeparameter infit.
- Returns:
self – The updated object.
- Return type:
object
- set_score_request(*, sample_weight='$UNCHANGED$')#
Configure whether metadata should be requested to be passed to the
scoremethod.Note that this method is only relevant when this estimator is used as a sub-estimator within a meta-estimator and metadata routing is enabled with
enable_metadata_routing=True(seesklearn.set_config()). Please check the User Guide on how the routing mechanism works.The options for each parameter are:
True: metadata is requested, and passed toscoreif provided. The request is ignored if metadata is not provided.False: metadata is not requested and the meta-estimator will not pass it toscore.None: metadata is not requested, and the meta-estimator will raise an error if the user provides it.str: metadata should be passed to the meta-estimator with this given alias instead of the original name.
The default (
sklearn.utils.metadata_routing.UNCHANGED) retains the existing request. This allows you to change the request for some parameters and not others.Added in version 1.3.
- Parameters:
sample_weight (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for
sample_weightparameter inscore.- Returns:
self – The updated object.
- Return type:
object
RuleFineTuneClassifier#
- class ex_fuzzy.classifiers.RuleFineTuneClassifier(nRules=30, nAnts=4, fuzzy_type=FUZZY_SETS.t1, tolerance=0.0, verbose=False, n_class=None, runner=1, expansion_factor=1, linguistic_variables=None, n_gen=30, pop_size=50, patience=10, random_state=33)[source]#
Bases:
ClassifierMixin,BaseEstimatorA classifier that works by mining a set of candidate rules with a minimum support and then uses a two step genetic optimization that chooses the optimal combination of those rules and fine tunes them.
- __init__(nRules=30, nAnts=4, fuzzy_type=FUZZY_SETS.t1, tolerance=0.0, verbose=False, n_class=None, runner=1, expansion_factor=1, linguistic_variables=None, n_gen=30, pop_size=50, patience=10, random_state=33)[source]#
Inits the optimizer with the corresponding parameters.
- Parameters:
nRules (int) – number of rules to optimize.
nAnts (int) – max number of antecedents to use.
fuzzy_type (FUZZY_SETS) – FUZZY_SET enum type in fuzzy_sets module. The kind of fuzzy set used.
tolerance (float) – tolerance for the dominance score of the rules.
verbose – if True, prints the progress of the optimization.
n_class (int) – number of classes in the problem. If None (default) the classifier will compute it empirically.
linguistic_variables (list[fuzzyVariable]) – linguistic variables per antecedent.
n_gen (int) – number of generations of the genetic search. fit can override it for one call.
pop_size (int) – population size of the genetic search. fit can override it for one call.
patience (int | None) – generations without improvement before the search stops early; None runs every generation.
random_state (int) – random seed of the genetic search.
- property fl_classifier1: BaseFuzzyRulesClassifier#
The classifier of the first phase, which selects among the mined rules. Built when fit is called.
- property fl_classifier2: BaseFuzzyRulesClassifier#
The classifier of the second phase, which fine tunes the selected rules. Built when fit is called.
- fit(X, y, n_gen=CONSTRUCTOR, pop_size=CONSTRUCTOR, checkpoints=0, **kwargs)[source]#
Trains the model with the given data.
- Parameters:
X (array) – samples to train.
y (array) – labels for each sample.
n_gen (int) – number of generations to compute in the genetic optimization. Defaults to the constructor’s value.
pop_size (int) – number of subjects per generation. Defaults to the constructor’s value.
checkpoints (int) – if bigger than 0, will save the best subject per x generations in a text file.
kwargs – additional parameters for the genetic optimization, including early stopping with patience=10 and min_delta=1e-4 by default. See fit method in BaseRuleBaseClassifier.
- Returns:
the fitted classifier.
- predict(X)[source]#
Predict for each sample the corresponding class.
- Parameters:
X (array) – samples to predict.
- Returns:
a class for each sample.
- Return type:
array
- set_fit_request(*, checkpoints='$UNCHANGED$', n_gen='$UNCHANGED$', pop_size='$UNCHANGED$')#
Configure whether metadata should be requested to be passed to the
fitmethod.Note that this method is only relevant when this estimator is used as a sub-estimator within a meta-estimator and metadata routing is enabled with
enable_metadata_routing=True(seesklearn.set_config()). Please check the User Guide on how the routing mechanism works.The options for each parameter are:
True: metadata is requested, and passed tofitif provided. The request is ignored if metadata is not provided.False: metadata is not requested and the meta-estimator will not pass it tofit.None: metadata is not requested, and the meta-estimator will raise an error if the user provides it.str: metadata should be passed to the meta-estimator with this given alias instead of the original name.
The default (
sklearn.utils.metadata_routing.UNCHANGED) retains the existing request. This allows you to change the request for some parameters and not others.Added in version 1.3.
- Parameters:
checkpoints (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for
checkpointsparameter infit.n_gen (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for
n_genparameter infit.pop_size (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for
pop_sizeparameter infit.
- Returns:
self – The updated object.
- Return type:
object
- set_score_request(*, sample_weight='$UNCHANGED$')#
Configure whether metadata should be requested to be passed to the
scoremethod.Note that this method is only relevant when this estimator is used as a sub-estimator within a meta-estimator and metadata routing is enabled with
enable_metadata_routing=True(seesklearn.set_config()). Please check the User Guide on how the routing mechanism works.The options for each parameter are:
True: metadata is requested, and passed toscoreif provided. The request is ignored if metadata is not provided.False: metadata is not requested and the meta-estimator will not pass it toscore.None: metadata is not requested, and the meta-estimator will raise an error if the user provides it.str: metadata should be passed to the meta-estimator with this given alias instead of the original name.
The default (
sklearn.utils.metadata_routing.UNCHANGED) retains the existing request. This allows you to change the request for some parameters and not others.Added in version 1.3.
- Parameters:
sample_weight (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for
sample_weightparameter inscore.- Returns:
self – The updated object.
- Return type:
object
Examples#
Basic Usage#
from ex_fuzzy.classifiers import RuleMineClassifier
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
# Load data
X, y = load_iris(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)
# Create and train classifier
classifier = RuleMineClassifier(nRules=20, nAnts=4, verbose=True)
classifier.fit(X_train, y_train)
# Make predictions
y_pred = classifier.predict(X_test)
accuracy = classifier.score(X_test, y_test)
print(f"Accuracy: {accuracy:.3f}")
See Also#
ex_fuzzy.evolutionary_fit: Underlying genetic optimizationex_fuzzy.rule_mining: Rule mining functionalityex_fuzzy.fuzzy_sets: Fuzzy set definitions