Evolutionary Fit Module#

The ex_fuzzy.evolutionary_fit module implements evolutionary optimization for fuzzy rule-based classifiers.

Overview#

This module provides:

Training performance#

The built-in T1/T2 classification objective automatically uses an optimized CPU evaluator. It decodes each chromosome with the reference rule constructor, computes firing strengths once, and reuses them for dominance scoring, pruning and final winner-rule prediction. Rules that never win a correctly classified sample are still removed. Reporting metrics are computed for the selected model rather than repeatedly for every candidate. Integer-label MCC uses a direct confusion-matrix calculation.

No public parameters change, and no compiler or additional dependency is needed. Fixed partitions retain their precomputed memberships; optimized partitions are recomputed for each chromosome. Custom losses, nonnumeric internal labels and other fuzzy types retain the full evaluator. EvoX classification also uses this CPU path. Regression and candidate-rule mining retain their existing evaluators.

To compare exact fitness values, selected chromosomes and predictions while measuring candidate evaluation and complete seeded fits, run from the repository root:

python benchmarks/benchmark_classifier_fitness.py --samples 1000

The script reports median times over three runs for both fixed and optimized partitions, using identical population sizes and generation budgets with early stopping disabled. Run it without competing workloads for useful timings. Speedups depend on data size, rule count and hardware. Population batching and additional parallelism are separate future optimization steps.

BaseFuzzyRulesClassifier#

class ex_fuzzy.evolutionary_fit.BaseFuzzyRulesClassifier(nRules=30, nAnts=4, fuzzy_type=FUZZY_SETS.t1, tolerance=0.0, class_names=None, n_linguistic_variables=3, verbose=False, linguistic_variables=None, categorical_mask=None, domain=None, n_class=None, precomputed_rules=None, runner=1, ds_mode=0, allow_unknown=False, backend='pymoo', detect_categorical=True, n_gen=70, pop_size=30, patience=10, min_delta=0.0001, random_state=33, var_prob=0.3, sbx_eta=3.0, mutation_eta=7.0, tournament_size=3)[source]#

Bases: ClassifierMixin, BaseEstimator

Class that is used as a classifier for a fuzzy rule based system. Supports precomputed and optimization of the linguistic variables.

__init__(nRules=30, nAnts=4, fuzzy_type=FUZZY_SETS.t1, tolerance=0.0, class_names=None, n_linguistic_variables=3, verbose=False, linguistic_variables=None, categorical_mask=None, domain=None, n_class=None, precomputed_rules=None, runner=1, ds_mode=0, allow_unknown=False, backend='pymoo', detect_categorical=True, n_gen=70, pop_size=30, patience=10, min_delta=0.0001, random_state=33, var_prob=0.3, sbx_eta=3.0, mutation_eta=7.0, tournament_size=3)[source]#

Inits the optimizer with the corresponding parameters.

Parameters:
  • nRules (int) – number of rules to optimize.

  • nAnts (int) – max number of antecedents to use.

  • fuzzy_type (FUZZY_SETS) – FUZZY_SET enum type in fuzzy_sets module. The kind of fuzzy set used.

  • tolerance (float) – tolerance for the dominance score of the rules.

  • n_linguist_variables – number of linguistic variables per antecedent.

  • verbose – if True, prints the progress of the optimization.

  • linguistic_variables (list[fuzzyVariable]) – list of fuzzyVariables type. If None (default) the optimization process will init+optimize them.

  • domain (list[float]) – list of the limits for each variable. If None (default) the classifier will compute them empirically.

  • n_class (int) – number of classes in the problem. If None (default) the classifier will compute it empirically.

  • precomputed_rules (MasterRuleBase) – MasterRuleBase object. If not None, the classifier will use the rules in the object and ignore the conflicting parameters.

  • runner (int) – number of threads used to evaluate candidates. Threads disable the fit-local fitness and firing caches, so a serial fit (1, the default) is usually faster; more threads only pay off for an expensive custom loss.

  • ds_mode (int | str) – inference weighting mode: 0 or ‘dominance’ weights rules by their dominance score, 1 or ‘unweighted’ uses the firing strengths alone, 2 or ‘optimized’ lets the genetic search set a weight per rule.

  • allow_unknown (bool) – if True, the classifier will allow the unknown class in the classification process. (Which would be a -1 value)

  • backend (str) – evolutionary backend to use. Options: ‘pymoo’ (default, CPU) or ‘evox’ (GPU-accelerated). Install with: pip install ex-fuzzy[evox]

  • detect_categorical (bool) – if True (default) and no categorical_mask is given, the categorical variables are detected from the data with utils.detect_categorical_mask. Ignored when categorical_mask is given or when linguistic_variables are precomputed.

  • n_gen (int) – number of generations of the genetic search. fit can override it for one call.

  • pop_size (int) – population size of the genetic search. fit can override it for one call.

  • patience (int | None) – generations without improvement before the search stops early; None runs every generation.

  • min_delta (float) – minimum fitness improvement that resets the patience.

  • random_state (int) – random seed of the genetic search.

  • var_prob (float) – crossover probability.

  • sbx_eta (float) – eta parameter of the SBX crossover.

  • mutation_eta (float) – eta parameter of the polynomial mutation.

  • tournament_size (int) – size of the selection tournament.

customized_loss(loss_function)[source]#

Function to customize the loss function used for the optimization.

Parameters:

loss_function – function that takes as input the true labels and the predicted labels and returns a float.

Returns:

None

fit(X, y, n_gen=CONSTRUCTOR, pop_size=CONSTRUCTOR, checkpoints=0, candidate_rules=None, initial_rules=None, random_state=CONSTRUCTOR, var_prob=CONSTRUCTOR, sbx_eta=CONSTRUCTOR, mutation_eta=CONSTRUCTOR, tournament_size=CONSTRUCTOR, bootstrap_size=1000, checkpoint_path='', p_value_compute=False, checkpoint_callback=None, patience=CONSTRUCTOR, min_delta=CONSTRUCTOR)[source]#

Fits a fuzzy rule based classifier using a genetic algorithm to the given data.

The search settings (n_gen, pop_size, patience, min_delta, random_state, var_prob, sbx_eta, mutation_eta and tournament_size) default to the values given to the constructor; passing them here overrides them for this fit only.

Parameters:
  • X (array) – numpy array samples x features

  • y (array) – labels. integer array samples (x 1)

  • n_gen (int) – integer. Number of generations to run the genetic algorithm.

  • pop_size (int) – integer. Population size for each gneration.

  • checkpoints (int) – integer. Number of checkpoints to save the best rulebase found so far.

  • candidate_rules (MasterRuleBase) – if these rules exist, the optimization process will choose the best rules from this set. If None (default) the rules will be generated from scratch.

  • initial_rules (MasterRuleBase) – if these rules exist, the optimization process will start from this set. If None (default) the rules will be generated from scratch.

  • random_state (int) – integer. Random seed for the optimization process.

  • var_prob (float) – float. Probability of crossover for the genetic algorithm.

  • sbx_eta (float) – float. Eta parameter for the SBX crossover.

  • checkpoint_path (str) – string. Path to save the checkpoints. If None (default) the checkpoints will be saved in the current directory.

  • mutation_eta (float) – float. Eta parameter for the polynomial mutation.

  • tournament_size (int) – integer. Size of the tournament for the genetic algorithm.

  • checkpoint_callback (Callable[[int, MasterRuleBase], None]) – function. Callback function that get executed at each checkpoint (‘checkpoints’ must be greater than 0), its arguments are the generation number and the rule_base of the checkpoint.

  • patience (int | None) – integer. Stop early when the best fitness does not improve for this many generations. Default is 10. Use None to run all generations.

  • min_delta (float) – float. Minimum fitness improvement required to reset patience. Default is 1e-4.

Returns:

the fitted classifier.

print_rule_bootstrap_results()[source]#

Prints the bootstrap results for each rule.

p_value_validation(bootstrap_size=100)[source]#

Computes the permutation and bootstrapping p-values for the classifier and its rules.

Parameters:

bootstrap_size (int) – integer. Number of bootstraps samples to use.

load_master_rule_base(rule_base)[source]#

Loads a master rule base to be used in the prediction process.

Parameters:

rule_base (MasterRuleBase) – ruleBase object.

Returns:

None

Return type:

None

explainable_predict(X, out_class_names=False)[source]#

Returns the predicted class for each sample, with the winning rule, its association degree and confidence interval.

Parameters:
  • X (array) – np array samples x features.

  • out_class_names – if True, the predictions are the consequent names instead of the fitted labels.

Returns:

predictions, winning rules, winning association degrees and confidence intervals.

Return type:

an ExplainedPrediction named tuple

forward(X, out_class_names=False)[source]#

Returns the predicted class for each sample.

Parameters:
  • X (array) – np array samples x features.

  • out_class_names – if True, the output will be the consequent names instead of the fitted labels.

Returns:

np array samples (x 1) with the predicted class.

Return type:

array

predict(X, out_class_names=False)[source]#

Returns the predicted class for each sample.

A fitted classifier predicts the labels it was fitted with (see classes_); samples where no rule fires are -1 for numeric labels and ‘Unknown’ otherwise. A classifier built from precomputed rules predicts consequent indexes.

Parameters:
  • X (array) – np array samples x features.

  • out_class_names – if True, the output will be the consequent names instead of the fitted labels.

Returns:

np array samples (x 1) with the predicted class.

Return type:

array

predict_proba_rules(X, truth_degrees=True)[source]#

Returns the predicted class probabilities for each sample.

Parameters:
  • X (array) – np array samples x features.

  • truth_degrees (bool) – if True, the output will be the truth degrees of the rules. If false, will return the association degrees i.e. the truth degree multiplied by the weights/dominance of the rules. (depending on the inference mode chosen)

Returns:

np array samples x classes with the predicted class probabilities.

Return type:

array

predict_membership_class(X)[source]#

Returns the predicted class memberships for each sample.

Parameters:

X (array) – np array samples x features.

Returns:

np array samples x classes with the predicted class probabilities.

Return type:

array

predict_proba(X)[source]#

Returns the predicted class probabilities for each sample.

Parameters:

X (array) – np array samples x features.

Returns:

np array samples x classes with the predicted class probabilities.

Return type:

array

print_rules(return_rules=False, bootstrap_results=False)[source]#

Print the rules contained in the fitted rulebase.

plot_fuzzy_variables()[source]#

Plot the fuzzy partitions in each fuzzy variable.

rename_fuzzy_variables()[source]#

Renames the linguist labels so that high, low and so on are consistent. It does so usually after an optimization process.

Returns:

None. Names are sorted accorded to the central point of the fuzzy memberships.

Return type:

None

get_rulebase()[source]#

Get the rulebase obtained after fitting the classifier to the data.

Returns:

a matrix format for the rulebase.

Return type:

list[array]

reparametrize_loss(alpha, beta)[source]#

Changes the parameters in the loss function.

Parameters:
  • alpha (float) – controls the MCC term.

  • beta (float) – controls the average rule size loss.

Note

Does not check for convexity preservation. The user can play with these parameters as it wills.

reparametrice_loss(alpha, beta)[source]#

Deprecated spelling of reparametrize_loss.

__call__(X)[source]#

Returns the predicted class for each sample.

Parameters:

X (array) – np array samples x features.

Returns:

np array samples (x 1) with the predicted class.

Return type:

array

set_fit_request(*, bootstrap_size='$UNCHANGED$', candidate_rules='$UNCHANGED$', checkpoint_callback='$UNCHANGED$', checkpoint_path='$UNCHANGED$', checkpoints='$UNCHANGED$', initial_rules='$UNCHANGED$', min_delta='$UNCHANGED$', mutation_eta='$UNCHANGED$', n_gen='$UNCHANGED$', p_value_compute='$UNCHANGED$', patience='$UNCHANGED$', pop_size='$UNCHANGED$', random_state='$UNCHANGED$', sbx_eta='$UNCHANGED$', tournament_size='$UNCHANGED$', var_prob='$UNCHANGED$')#

Configure whether metadata should be requested to be passed to the fit method.

Note that this method is only relevant when this estimator is used as a sub-estimator within a meta-estimator and metadata routing is enabled with enable_metadata_routing=True (see sklearn.set_config()). Please check the User Guide on how the routing mechanism works.

The options for each parameter are:

  • True: metadata is requested, and passed to fit if provided. The request is ignored if metadata is not provided.

  • False: metadata is not requested and the meta-estimator will not pass it to fit.

  • None: metadata is not requested, and the meta-estimator will raise an error if the user provides it.

  • str: metadata should be passed to the meta-estimator with this given alias instead of the original name.

The default (sklearn.utils.metadata_routing.UNCHANGED) retains the existing request. This allows you to change the request for some parameters and not others.

Added in version 1.3.

Parameters:
  • bootstrap_size (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for bootstrap_size parameter in fit.

  • candidate_rules (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for candidate_rules parameter in fit.

  • checkpoint_callback (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for checkpoint_callback parameter in fit.

  • checkpoint_path (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for checkpoint_path parameter in fit.

  • checkpoints (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for checkpoints parameter in fit.

  • initial_rules (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for initial_rules parameter in fit.

  • min_delta (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for min_delta parameter in fit.

  • mutation_eta (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for mutation_eta parameter in fit.

  • n_gen (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for n_gen parameter in fit.

  • p_value_compute (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for p_value_compute parameter in fit.

  • patience (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for patience parameter in fit.

  • pop_size (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for pop_size parameter in fit.

  • random_state (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for random_state parameter in fit.

  • sbx_eta (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for sbx_eta parameter in fit.

  • tournament_size (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for tournament_size parameter in fit.

  • var_prob (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for var_prob parameter in fit.

Returns:

self – The updated object.

Return type:

object

set_predict_request(*, out_class_names='$UNCHANGED$')#

Configure whether metadata should be requested to be passed to the predict method.

Note that this method is only relevant when this estimator is used as a sub-estimator within a meta-estimator and metadata routing is enabled with enable_metadata_routing=True (see sklearn.set_config()). Please check the User Guide on how the routing mechanism works.

The options for each parameter are:

  • True: metadata is requested, and passed to predict if provided. The request is ignored if metadata is not provided.

  • False: metadata is not requested and the meta-estimator will not pass it to predict.

  • None: metadata is not requested, and the meta-estimator will raise an error if the user provides it.

  • str: metadata should be passed to the meta-estimator with this given alias instead of the original name.

The default (sklearn.utils.metadata_routing.UNCHANGED) retains the existing request. This allows you to change the request for some parameters and not others.

Added in version 1.3.

Parameters:

out_class_names (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for out_class_names parameter in predict.

Returns:

self – The updated object.

Return type:

object

set_score_request(*, sample_weight='$UNCHANGED$')#

Configure whether metadata should be requested to be passed to the score method.

Note that this method is only relevant when this estimator is used as a sub-estimator within a meta-estimator and metadata routing is enabled with enable_metadata_routing=True (see sklearn.set_config()). Please check the User Guide on how the routing mechanism works.

The options for each parameter are:

  • True: metadata is requested, and passed to score if provided. The request is ignored if metadata is not provided.

  • False: metadata is not requested and the meta-estimator will not pass it to score.

  • None: metadata is not requested, and the meta-estimator will raise an error if the user provides it.

  • str: metadata should be passed to the meta-estimator with this given alias instead of the original name.

The default (sklearn.utils.metadata_routing.UNCHANGED) retains the existing request. This allows you to change the request for some parameters and not others.

Added in version 1.3.

Parameters:

sample_weight (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for sample_weight parameter in score.

Returns:

self – The updated object.

Return type:

object

FitRuleBase#

class ex_fuzzy.evolutionary_fit.FitRuleBase(X, y, nRules, nAnts, n_classes, thread_runner=None, linguistic_variables=None, n_linguistic_variables=3, fuzzy_type=FUZZY_SETS.t1, domain=None, categorical_mask=None, tolerance=0.01, alpha=0.0, beta=0.0, ds_mode=0, allow_unknown=False, backend_name='pymoo', var_names=None)[source]#

Bases: Problem

Class to model, independently of the optimizer, the fitting of a rulebase for a classification problem using Evolutionary strategies. Supports type 1 and iv fs (iv-type 2)

vl_names = [[], [], ['Low', 'High'], ['Low', 'Medium', 'High'], ['Low', 'Medium', 'High', 'Very High'], ['Very Low', 'Low', 'Medium', 'High', 'Very High']]#
__init__(X, y, nRules, nAnts, n_classes, thread_runner=None, linguistic_variables=None, n_linguistic_variables=3, fuzzy_type=FUZZY_SETS.t1, domain=None, categorical_mask=None, tolerance=0.01, alpha=0.0, beta=0.0, ds_mode=0, allow_unknown=False, backend_name='pymoo', var_names=None)[source]#

Cosntructor method. Initializes the classifier with the number of antecedents, linguist variables and the kind of fuzzy set desired.

Parameters:
  • X (array) – np array or pandas dataframe samples x features.

  • y (array) – np vector containing the target classes. vector sample

  • nRules (int) – number of rules to optimize.

  • nAnts (int) – max number of antecedents to use.

  • n_class – number of classes in the problem. If None (as default) it will be computed from the data.

  • linguistic_variables (list[fuzzyVariable]) – list of linguistic variables precomputed. If given, the rest of conflicting arguments are ignored.

  • n_linguistic_variables (int) – number of linguistic variables per antecedent.

  • fuzzy_type – Define the fuzzy set or fuzzy set extension used as linguistic variable.

  • domain (list) – list with the upper and lower domains of each input variable. If None (as default) it will stablish the empirical min/max as the limits.

  • tolerance (float) – float. Tolerance for the size evaluation.

  • alpha (float) – float. Weight for the rulebase size term in the fitness function. (Penalizes number of rules)

  • beta (float) – float. Weight for the average rule size term in the fitness function.

  • ds_mode (int) – int. Mode for the dominance score. 0: normal dominance score, 1: rules without weights, 2: weights optimized for each rule based on the data.

  • allow_unknown (bool) – if True, the classifier will allow the unknown class in the classification process. (Which would be a -1 value)

  • var_names (list) – list of variable names. If None, extracted from DataFrame columns or auto-generated.

encode_rulebase(rule_base, optimize_lv)[source]#

Given a rule base, constructs the corresponding gene associated with that rule base.

GENE STRUCTURE

First: antecedents chosen by each rule. Size: nAnts * nRules (index of the antecedent) Second: Variable linguistics used. Size: nAnts * nRules Third: Parameters for the fuzzy partitions of the chosen variables. Size: nAnts * self.n_linguistic_variables * 8|4 (2 trapezoidal memberships if t2) Four: Consequent classes. Size: nRules

Parameters:
  • rule_base (MasterRuleBase) – rule base object.

  • optimize_lv (bool) – must be False: only genes over fixed linguistic variables can be encoded.

Returns:

np array of size self.single_gen_size.

Raises:
  • NotImplementedError – if optimize_lv is True.

  • ValueError – if the problem does not have one antecedent slot per feature.

Return type:

array

array_evaluation = True#

Set to False to force the object decoder, for benchmarks and parity tests.

torch_devices = ('cuda',)#

Device types on which populations may be scored by the exact PyTorch objective. Scoring CPU tensors gains nothing over the NumPy routes, so only CUDA is enabled; parity tests add ‘cpu’ to run the same code.

fitness_func(ruleBase, X, y, tolerance, alpha=0.0, beta=0.0, precomputed_truth=None)[source]#

Fitness function for the optimization problem.

Parameters:
  • ruleBase (RuleBase) – RuleBase object

  • X (array) – array of train samples. X shape = (n_samples, n_features)

  • y (array) – array of train labels. y shape = (n_samples,)

  • tolerance (float) – float. Tolerance for the size evaluation.

  • alpha (float) – float. Weight for the accuracy term.

  • beta (float) – float. Weight for the average rule size term.

  • precomputed_truth (array) – np array. If given, it will be used as the truth values for the evaluation.

Returns:

float. Fitness value.

Return type:

float

See Also#

  • ex_fuzzy.classifiers

  • ex_fuzzy.rules

  • ex_fuzzy.rule_mining

  • ex_fuzzy.eval_tools