Classifiers Module#

The ex_fuzzy.classifiers module provides the main classification interface for the ex-fuzzy library.

Overview#

This module contains high-level classifiers built on fuzzy rule mining. FuzzyRulesClassifier mines fuzzy association rules on a capped feature space and selects a compact rule base with a genetic algorithm (see Fuzzy association rule classifier). RuleMineClassifier and RuleFineTuneClassifier pass mined candidate rules to BaseFuzzyRulesClassifier.

Classes#

FuzzyRulesClassifier#

class ex_fuzzy.classifiers.FuzzyRulesClassifier(nRules=None, nAnts=3, fuzzy_type=FUZZY_SETS.t1, tolerance=0.0, verbose=False, n_class=None, runner=1, expansion_factor=1, linguistic_variables=None, max_features=8, feature_selector='mutual_info', feature_selection='per_class', n_linguistic_variables='auto', min_support=0.05, min_confidence=0.5, candidates_per_class=50, rule_mode='additive', n_gen=100, pop_size=60, patience=20, rule_penalty=0.001, rules_per_class=None, random_state=None)[source]#

Bases: ClassifierMixin, BaseEstimator

Fuzzy association rule classifier with feature capping and genetic rule selection.

The classifier follows the FARC-HD family of fuzzy association rule classifiers (Alcalá-Fdez, Alcalá and Herrera, IEEE Transactions on Fuzzy Systems 19(5), 2011), which learns a compact linguistic rule base:

  1. The input space is capped to the max_features most informative features, which keeps rule generation tractable and removes noise.

  2. Each kept feature gets a fixed fuzzy partition (quantile-based terms for numerical features, one term per category for categorical ones).

  3. Candidate rules with up to nAnts conditions are mined per class and filtered by support, confidence and penalized certainty factor.

  4. A covering-based subgroup-discovery prescreen keeps a diverse pool of up to candidates_per_class rules per class.

  5. A genetic algorithm selects a compact subset of rules that maximizes training accuracy with a small penalty per rule, capped by nRules or rules_per_class when set.

Rules are weighted by their penalized certainty factor and combined as in Ex-Fuzzy’s regression: with rule_mode="additive" every matching rule votes with its weighted firing, and with rule_mode="sufficient" only each sample’s strongest weighted rule decides. Samples that fire no rule take the training majority class. ex_fuzzy.BaseFuzzyRulesClassifier is not modified.

Parameters:
  • nRules (int or None, default=None) – Maximum number of rules in the final rule base. None removes the absolute cap.

  • nAnts (int or "auto", default=3) – Maximum number of conditions per rule. "auto" allows four conditions when a class’s rules are mined from at most six features, where interactions must carry the model, and three otherwise. Four or five conditions are supported; the candidate search is then pruned by min_support, so raising it also speeds up long rules.

  • fuzzy_type (FUZZY_SETS, default=FUZZY_SETS.t1) – Only Type-1 fuzzy sets are supported.

  • tolerance (float, default=0.0) – Minimum penalized certainty factor a candidate rule must exceed.

  • verbose (bool, default=False) – Print a summary of the fitting stages.

  • n_class (int, optional) – Ignored; the classes are read from y. Kept for compatibility.

  • runner (int, default=1) – Ignored; the rule selection is vectorized. Kept for compatibility.

  • expansion_factor (int, default=1) – Multiplies candidates_per_class, enlarging the pool the genetic search chooses from.

  • linguistic_variables (list of fuzzyVariable, optional) – One fuzzy variable per input feature. When omitted, partitions are built from the data.

  • max_features (int, "auto" or None, default=8) – Maximum number of input features to keep. None keeps them all. "auto" keeps 8 for problems with up to five classes and 16 for problems with more, which need more features to separate.

  • feature_selector ({"mutual_info", "f_classif"} or callable, default="mutual_info") – Feature relevance score. A callable receives (X, y) and returns one score per feature.

  • feature_selection ({"global", "per_class"}, default="per_class") – "global" keeps the max_features most relevant features for all classes. "per_class" scores features one class against the rest and mines each class’s rules on its own max_features features, so the rule base can use more features in total while each rule search stays small.

  • n_linguistic_variables (int or "auto", default="auto") – Number of fuzzy terms per numerical feature. "auto" uses three terms for binary problems and five otherwise.

  • min_support (float, default=0.05) – Minimum class support of a candidate rule.

  • min_confidence (float, default=0.5) – Minimum confidence of a candidate rule.

  • candidates_per_class (int, default=50) – Maximum rules per class kept by the prescreen.

  • rule_mode ({"additive", "sufficient"}, default="additive") – How rules are combined into a class decision. "additive" sums each class’s weighted rule firing; "sufficient" keeps only each sample’s strongest weighted rule, as in ex_fuzzy.BaseFuzzyRulesRegressor.

  • n_gen (int, default=100) – Maximum generations of the rule-selection search.

  • pop_size (int, default=60) – Population size of the rule-selection search.

  • patience (int, default=20) – Generations without improvement before the search stops.

  • rule_penalty (float, default=1e-3) – Accuracy traded for each selected rule per class.

  • rules_per_class (int, optional) – Cap the rule base at rules_per_class times the number of classes. When nRules is also set, the smaller cap applies.

  • random_state (int, optional) – Seed for feature scoring and the genetic search.

classes_#

Class labels.

Type:

np.ndarray

selected_features_#

Indices of the input features the rules use.

Type:

np.ndarray

class_features_#

Input feature indices each class’s rules were mined from.

Type:

list of np.ndarray

max_features_#

Resolved feature cap.

Type:

int or None

n_conditions_#

Resolved maximum number of conditions per class.

Type:

list of int

linguistic_variables_#

Partitions of the selected features.

Type:

list of fuzzyVariable

rule_base_#

Selected rules, one rule base per class, weighted by certainty factor.

Type:

MasterRuleBase or None

n_rules_#

Number of selected rules.

Type:

int

majority_class_#

Index of the class predicted when no rule fires.

Type:

int

__init__(nRules=None, nAnts=3, fuzzy_type=FUZZY_SETS.t1, tolerance=0.0, verbose=False, n_class=None, runner=1, expansion_factor=1, linguistic_variables=None, max_features=8, feature_selector='mutual_info', feature_selection='per_class', n_linguistic_variables='auto', min_support=0.05, min_confidence=0.5, candidates_per_class=50, rule_mode='additive', n_gen=100, pop_size=60, patience=20, rule_penalty=0.001, rules_per_class=None, random_state=None)[source]#
fit(X, y, n_gen=None, pop_size=None, checkpoints=0, **kwargs)[source]#

Mine, prescreen and select the rule base.

Parameters:
  • X (array-like of shape (n_samples, n_features)) – Training features.

  • y (array-like of shape (n_samples,)) – Class labels.

  • n_gen (int) – int, optional: Override the constructor’s search budget for this fit.

  • pop_size (int) – int, optional: Override the constructor’s search budget for this fit.

  • checkpoints (int, default=0) – Ignored; kept for compatibility.

  • **kwargs – random_state and patience override the constructor values.

Returns:

FuzzyRulesClassifier

The fitted estimator.

predict(X)[source]#

Predict the class of each sample.

predict_proba(X)[source]#

Class probabilities from normalized rule scores.

Samples that fire no rule receive the training class distribution.

print_rules(return_rules=False)[source]#

Print the selected rules, one block per class.

internal_classifier()[source]#

The selected rule base wrapped as a BaseFuzzyRulesClassifier.

The wrapped model expects only the selected features (X[:, selected_features_]) and uses winning-rule inference without the majority-class fallback; use this estimator’s predict for the classifier’s own decisions.

set_fit_request(*, checkpoints='$UNCHANGED$', n_gen='$UNCHANGED$', pop_size='$UNCHANGED$')#

Configure whether metadata should be requested to be passed to the fit method.

Note that this method is only relevant when this estimator is used as a sub-estimator within a meta-estimator and metadata routing is enabled with enable_metadata_routing=True (see sklearn.set_config()). Please check the User Guide on how the routing mechanism works.

The options for each parameter are:

  • True: metadata is requested, and passed to fit if provided. The request is ignored if metadata is not provided.

  • False: metadata is not requested and the meta-estimator will not pass it to fit.

  • None: metadata is not requested, and the meta-estimator will raise an error if the user provides it.

  • str: metadata should be passed to the meta-estimator with this given alias instead of the original name.

The default (sklearn.utils.metadata_routing.UNCHANGED) retains the existing request. This allows you to change the request for some parameters and not others.

Added in version 1.3.

Parameters:
  • checkpoints (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for checkpoints parameter in fit.

  • n_gen (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for n_gen parameter in fit.

  • pop_size (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for pop_size parameter in fit.

Returns:

self – The updated object.

Return type:

object

set_score_request(*, sample_weight='$UNCHANGED$')#

Configure whether metadata should be requested to be passed to the score method.

Note that this method is only relevant when this estimator is used as a sub-estimator within a meta-estimator and metadata routing is enabled with enable_metadata_routing=True (see sklearn.set_config()). Please check the User Guide on how the routing mechanism works.

The options for each parameter are:

  • True: metadata is requested, and passed to score if provided. The request is ignored if metadata is not provided.

  • False: metadata is not requested and the meta-estimator will not pass it to score.

  • None: metadata is not requested, and the meta-estimator will raise an error if the user provides it.

  • str: metadata should be passed to the meta-estimator with this given alias instead of the original name.

The default (sklearn.utils.metadata_routing.UNCHANGED) retains the existing request. This allows you to change the request for some parameters and not others.

Added in version 1.3.

Parameters:

sample_weight (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for sample_weight parameter in score.

Returns:

self – The updated object.

Return type:

object

RuleMineClassifier#

class ex_fuzzy.classifiers.RuleMineClassifier(nRules=30, nAnts=4, fuzzy_type=FUZZY_SETS.t1, tolerance=0.0, verbose=False, n_class=None, runner=1, linguistic_variables=None, n_gen=30, pop_size=50, patience=10, random_state=33)[source]#

Bases: ClassifierMixin, BaseEstimator

A classifier that works by mining a set of candidate rules with a minimum support, confidence and lift, and then using a genetic algorithm that chooses the optimal combination of those rules.

The main classifier that mines candidate rules and then optimizes them using genetic algorithms.

__init__(nRules=30, nAnts=4, fuzzy_type=FUZZY_SETS.t1, tolerance=0.0, verbose=False, n_class=None, runner=1, linguistic_variables=None, n_gen=30, pop_size=50, patience=10, random_state=33)[source]#

Inits the optimizer with the corresponding parameters.

Parameters:
  • nRules (int) – number of rules to optimize.

  • nAnts (int) – max number of antecedents to use.

  • fuzzy_type (FUZZY_SETS) – FUZZY_SET enum type in fuzzy_sets module. The kind of fuzzy set used.

  • tolerance (float) – tolerance for the support/dominance score of the rules.

  • verbose – if True, prints the progress of the optimization.

  • n_class (int) – number of classes in the problem. If None (default) the classifier will compute it empirically.

  • runner (int) – number of threads to use.

  • linguistic_variables (list[fuzzyVariable]) – linguistic variables per antecedent.

  • n_gen (int) – number of generations of the genetic search. fit can override it for one call.

  • pop_size (int) – population size of the genetic search. fit can override it for one call.

  • patience (int | None) – generations without improvement before the search stops early; None runs every generation.

  • random_state (int) – random seed of the genetic search.

property fl_classifier: BaseFuzzyRulesClassifier#

The classifier that performs the final predictions. Built when fit is called.

fit(X, y, n_gen=CONSTRUCTOR, pop_size=CONSTRUCTOR, **kwargs)[source]#

Trains the model with the given data.

Parameters:
  • X (array) – samples to train.

  • y (array) – labels for each sample.

  • n_gen (int) – number of generations to compute in the genetic optimization. Defaults to the constructor’s value.

  • pop_size (int) – number of subjects per generation. Defaults to the constructor’s value.

  • kwargs – additional parameters for the genetic optimization, including early stopping with patience=10 and min_delta=1e-4 by default. See fit method in BaseRuleBaseClassifier.

Returns:

the fitted classifier.

predict(X)[source]#

Predict for each sample the corresponding class.

Parameters:

X (array) – samples to predict.

Returns:

a class for each sample.

Return type:

array

internal_classifier()[source]#
set_fit_request(*, n_gen='$UNCHANGED$', pop_size='$UNCHANGED$')#

Configure whether metadata should be requested to be passed to the fit method.

Note that this method is only relevant when this estimator is used as a sub-estimator within a meta-estimator and metadata routing is enabled with enable_metadata_routing=True (see sklearn.set_config()). Please check the User Guide on how the routing mechanism works.

The options for each parameter are:

  • True: metadata is requested, and passed to fit if provided. The request is ignored if metadata is not provided.

  • False: metadata is not requested and the meta-estimator will not pass it to fit.

  • None: metadata is not requested, and the meta-estimator will raise an error if the user provides it.

  • str: metadata should be passed to the meta-estimator with this given alias instead of the original name.

The default (sklearn.utils.metadata_routing.UNCHANGED) retains the existing request. This allows you to change the request for some parameters and not others.

Added in version 1.3.

Parameters:
  • n_gen (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for n_gen parameter in fit.

  • pop_size (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for pop_size parameter in fit.

Returns:

self – The updated object.

Return type:

object

set_score_request(*, sample_weight='$UNCHANGED$')#

Configure whether metadata should be requested to be passed to the score method.

Note that this method is only relevant when this estimator is used as a sub-estimator within a meta-estimator and metadata routing is enabled with enable_metadata_routing=True (see sklearn.set_config()). Please check the User Guide on how the routing mechanism works.

The options for each parameter are:

  • True: metadata is requested, and passed to score if provided. The request is ignored if metadata is not provided.

  • False: metadata is not requested and the meta-estimator will not pass it to score.

  • None: metadata is not requested, and the meta-estimator will raise an error if the user provides it.

  • str: metadata should be passed to the meta-estimator with this given alias instead of the original name.

The default (sklearn.utils.metadata_routing.UNCHANGED) retains the existing request. This allows you to change the request for some parameters and not others.

Added in version 1.3.

Parameters:

sample_weight (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for sample_weight parameter in score.

Returns:

self – The updated object.

Return type:

object

RuleFineTuneClassifier#

class ex_fuzzy.classifiers.RuleFineTuneClassifier(nRules=30, nAnts=4, fuzzy_type=FUZZY_SETS.t1, tolerance=0.0, verbose=False, n_class=None, runner=1, expansion_factor=1, linguistic_variables=None, n_gen=30, pop_size=50, patience=10, random_state=33)[source]#

Bases: ClassifierMixin, BaseEstimator

A classifier that works by mining a set of candidate rules with a minimum support and then uses a two step genetic optimization that chooses the optimal combination of those rules and fine tunes them.

__init__(nRules=30, nAnts=4, fuzzy_type=FUZZY_SETS.t1, tolerance=0.0, verbose=False, n_class=None, runner=1, expansion_factor=1, linguistic_variables=None, n_gen=30, pop_size=50, patience=10, random_state=33)[source]#

Inits the optimizer with the corresponding parameters.

Parameters:
  • nRules (int) – number of rules to optimize.

  • nAnts (int) – max number of antecedents to use.

  • fuzzy_type (FUZZY_SETS) – FUZZY_SET enum type in fuzzy_sets module. The kind of fuzzy set used.

  • tolerance (float) – tolerance for the dominance score of the rules.

  • verbose – if True, prints the progress of the optimization.

  • n_class (int) – number of classes in the problem. If None (default) the classifier will compute it empirically.

  • linguistic_variables (list[fuzzyVariable]) – linguistic variables per antecedent.

  • n_gen (int) – number of generations of the genetic search. fit can override it for one call.

  • pop_size (int) – population size of the genetic search. fit can override it for one call.

  • patience (int | None) – generations without improvement before the search stops early; None runs every generation.

  • random_state (int) – random seed of the genetic search.

property fl_classifier1: BaseFuzzyRulesClassifier#

The classifier of the first phase, which selects among the mined rules. Built when fit is called.

property fl_classifier2: BaseFuzzyRulesClassifier#

The classifier of the second phase, which fine tunes the selected rules. Built when fit is called.

fit(X, y, n_gen=CONSTRUCTOR, pop_size=CONSTRUCTOR, checkpoints=0, **kwargs)[source]#

Trains the model with the given data.

Parameters:
  • X (array) – samples to train.

  • y (array) – labels for each sample.

  • n_gen (int) – number of generations to compute in the genetic optimization. Defaults to the constructor’s value.

  • pop_size (int) – number of subjects per generation. Defaults to the constructor’s value.

  • checkpoints (int) – if bigger than 0, will save the best subject per x generations in a text file.

  • kwargs – additional parameters for the genetic optimization, including early stopping with patience=10 and min_delta=1e-4 by default. See fit method in BaseRuleBaseClassifier.

Returns:

the fitted classifier.

predict(X)[source]#

Predict for each sample the corresponding class.

Parameters:

X (array) – samples to predict.

Returns:

a class for each sample.

Return type:

array

internal_classifier()[source]#
set_fit_request(*, checkpoints='$UNCHANGED$', n_gen='$UNCHANGED$', pop_size='$UNCHANGED$')#

Configure whether metadata should be requested to be passed to the fit method.

Note that this method is only relevant when this estimator is used as a sub-estimator within a meta-estimator and metadata routing is enabled with enable_metadata_routing=True (see sklearn.set_config()). Please check the User Guide on how the routing mechanism works.

The options for each parameter are:

  • True: metadata is requested, and passed to fit if provided. The request is ignored if metadata is not provided.

  • False: metadata is not requested and the meta-estimator will not pass it to fit.

  • None: metadata is not requested, and the meta-estimator will raise an error if the user provides it.

  • str: metadata should be passed to the meta-estimator with this given alias instead of the original name.

The default (sklearn.utils.metadata_routing.UNCHANGED) retains the existing request. This allows you to change the request for some parameters and not others.

Added in version 1.3.

Parameters:
  • checkpoints (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for checkpoints parameter in fit.

  • n_gen (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for n_gen parameter in fit.

  • pop_size (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for pop_size parameter in fit.

Returns:

self – The updated object.

Return type:

object

set_score_request(*, sample_weight='$UNCHANGED$')#

Configure whether metadata should be requested to be passed to the score method.

Note that this method is only relevant when this estimator is used as a sub-estimator within a meta-estimator and metadata routing is enabled with enable_metadata_routing=True (see sklearn.set_config()). Please check the User Guide on how the routing mechanism works.

The options for each parameter are:

  • True: metadata is requested, and passed to score if provided. The request is ignored if metadata is not provided.

  • False: metadata is not requested and the meta-estimator will not pass it to score.

  • None: metadata is not requested, and the meta-estimator will raise an error if the user provides it.

  • str: metadata should be passed to the meta-estimator with this given alias instead of the original name.

The default (sklearn.utils.metadata_routing.UNCHANGED) retains the existing request. This allows you to change the request for some parameters and not others.

Added in version 1.3.

Parameters:

sample_weight (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for sample_weight parameter in score.

Returns:

self – The updated object.

Return type:

object

Examples#

Basic Usage#

from ex_fuzzy.classifiers import RuleMineClassifier
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split

# Load data
X, y = load_iris(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42)

# Create and train classifier
classifier = RuleMineClassifier(nRules=20, nAnts=4, verbose=True)
classifier.fit(X_train, y_train)

# Make predictions
y_pred = classifier.predict(X_test)
accuracy = classifier.score(X_test, y_test)
print(f"Accuracy: {accuracy:.3f}")

See Also#

  • ex_fuzzy.evolutionary_fit : Underlying genetic optimization

  • ex_fuzzy.rule_mining : Rule mining functionality

  • ex_fuzzy.fuzzy_sets : Fuzzy set definitions