Evolutionary Fit Module#
The ex_fuzzy.evolutionary_fit module implements evolutionary optimization
for fuzzy rule-based classifiers.
Overview#
This module provides:
BaseFuzzyRulesClassifierfor end-to-end fuzzy classifier training.FitRuleBaseas the optimization problem used internally.
Training performance#
The built-in T1/T2 classification objective automatically uses an optimized CPU evaluator. It decodes each chromosome with the reference rule constructor, computes firing strengths once, and reuses them for dominance scoring, pruning and final winner-rule prediction. Rules that never win a correctly classified sample are still removed. Reporting metrics are computed for the selected model rather than repeatedly for every candidate. Integer-label MCC uses a direct confusion-matrix calculation.
No public parameters change, and no compiler or additional dependency is needed. Fixed partitions retain their precomputed memberships; optimized partitions are recomputed for each chromosome. Custom losses, nonnumeric internal labels and other fuzzy types retain the full evaluator. EvoX classification also uses this CPU path. Regression and candidate-rule mining retain their existing evaluators.
To compare exact fitness values, selected chromosomes and predictions while measuring candidate evaluation and complete seeded fits, run from the repository root:
python benchmarks/benchmark_classifier_fitness.py --samples 1000
The script reports median times over three runs for both fixed and optimized partitions, using identical population sizes and generation budgets with early stopping disabled. Run it without competing workloads for useful timings. Speedups depend on data size, rule count and hardware. Population batching and additional parallelism are separate future optimization steps.
BaseFuzzyRulesClassifier#
- class ex_fuzzy.evolutionary_fit.BaseFuzzyRulesClassifier(nRules=30, nAnts=4, fuzzy_type=FUZZY_SETS.t1, tolerance=0.0, class_names=None, n_linguistic_variables=3, verbose=False, linguistic_variables=None, categorical_mask=None, domain=None, n_class=None, precomputed_rules=None, runner=1, ds_mode=0, allow_unknown=False, backend='pymoo', detect_categorical=True, n_gen=70, pop_size=30, patience=10, min_delta=0.0001, random_state=33, var_prob=0.3, sbx_eta=3.0, mutation_eta=7.0, tournament_size=3)[source]#
Bases:
ClassifierMixin,BaseEstimatorClass that is used as a classifier for a fuzzy rule based system. Supports precomputed and optimization of the linguistic variables.
- __init__(nRules=30, nAnts=4, fuzzy_type=FUZZY_SETS.t1, tolerance=0.0, class_names=None, n_linguistic_variables=3, verbose=False, linguistic_variables=None, categorical_mask=None, domain=None, n_class=None, precomputed_rules=None, runner=1, ds_mode=0, allow_unknown=False, backend='pymoo', detect_categorical=True, n_gen=70, pop_size=30, patience=10, min_delta=0.0001, random_state=33, var_prob=0.3, sbx_eta=3.0, mutation_eta=7.0, tournament_size=3)[source]#
Inits the optimizer with the corresponding parameters.
- Parameters:
nRules (int) – number of rules to optimize.
nAnts (int) – max number of antecedents to use.
fuzzy_type (FUZZY_SETS) – FUZZY_SET enum type in fuzzy_sets module. The kind of fuzzy set used.
tolerance (float) – tolerance for the dominance score of the rules.
n_linguist_variables – number of linguistic variables per antecedent.
verbose – if True, prints the progress of the optimization.
linguistic_variables (list[fuzzyVariable]) – list of fuzzyVariables type. If None (default) the optimization process will init+optimize them.
domain (list[float]) – list of the limits for each variable. If None (default) the classifier will compute them empirically.
n_class (int) – number of classes in the problem. If None (default) the classifier will compute it empirically.
precomputed_rules (MasterRuleBase) – MasterRuleBase object. If not None, the classifier will use the rules in the object and ignore the conflicting parameters.
runner (int) – number of threads used to evaluate candidates. Threads disable the fit-local fitness and firing caches, so a serial fit (1, the default) is usually faster; more threads only pay off for an expensive custom loss.
ds_mode (int | str) – inference weighting mode: 0 or ‘dominance’ weights rules by their dominance score, 1 or ‘unweighted’ uses the firing strengths alone, 2 or ‘optimized’ lets the genetic search set a weight per rule.
allow_unknown (bool) – if True, the classifier will allow the unknown class in the classification process. (Which would be a -1 value)
backend (str) – evolutionary backend to use. Options: ‘pymoo’ (default, CPU) or ‘evox’ (GPU-accelerated). Install with: pip install ex-fuzzy[evox]
detect_categorical (bool) – if True (default) and no categorical_mask is given, the categorical variables are detected from the data with utils.detect_categorical_mask. Ignored when categorical_mask is given or when linguistic_variables are precomputed.
n_gen (int) – number of generations of the genetic search. fit can override it for one call.
pop_size (int) – population size of the genetic search. fit can override it for one call.
patience (int | None) – generations without improvement before the search stops early; None runs every generation.
min_delta (float) – minimum fitness improvement that resets the patience.
random_state (int) – random seed of the genetic search.
var_prob (float) – crossover probability.
sbx_eta (float) – eta parameter of the SBX crossover.
mutation_eta (float) – eta parameter of the polynomial mutation.
tournament_size (int) – size of the selection tournament.
- customized_loss(loss_function)[source]#
Function to customize the loss function used for the optimization.
- Parameters:
loss_function – function that takes as input the true labels and the predicted labels and returns a float.
- Returns:
None
- fit(X, y, n_gen=CONSTRUCTOR, pop_size=CONSTRUCTOR, checkpoints=0, candidate_rules=None, initial_rules=None, random_state=CONSTRUCTOR, var_prob=CONSTRUCTOR, sbx_eta=CONSTRUCTOR, mutation_eta=CONSTRUCTOR, tournament_size=CONSTRUCTOR, bootstrap_size=1000, checkpoint_path='', p_value_compute=False, checkpoint_callback=None, patience=CONSTRUCTOR, min_delta=CONSTRUCTOR)[source]#
Fits a fuzzy rule based classifier using a genetic algorithm to the given data.
The search settings (n_gen, pop_size, patience, min_delta, random_state, var_prob, sbx_eta, mutation_eta and tournament_size) default to the values given to the constructor; passing them here overrides them for this fit only.
- Parameters:
X (array) – numpy array samples x features
y (array) – labels. integer array samples (x 1)
n_gen (int) – integer. Number of generations to run the genetic algorithm.
pop_size (int) – integer. Population size for each gneration.
checkpoints (int) – integer. Number of checkpoints to save the best rulebase found so far.
candidate_rules (MasterRuleBase) – if these rules exist, the optimization process will choose the best rules from this set. If None (default) the rules will be generated from scratch.
initial_rules (MasterRuleBase) – if these rules exist, the optimization process will start from this set. If None (default) the rules will be generated from scratch.
random_state (int) – integer. Random seed for the optimization process.
var_prob (float) – float. Probability of crossover for the genetic algorithm.
sbx_eta (float) – float. Eta parameter for the SBX crossover.
checkpoint_path (str) – string. Path to save the checkpoints. If None (default) the checkpoints will be saved in the current directory.
mutation_eta (float) – float. Eta parameter for the polynomial mutation.
tournament_size (int) – integer. Size of the tournament for the genetic algorithm.
checkpoint_callback (Callable[[int, MasterRuleBase], None]) – function. Callback function that get executed at each checkpoint (‘checkpoints’ must be greater than 0), its arguments are the generation number and the rule_base of the checkpoint.
patience (int | None) – integer. Stop early when the best fitness does not improve for this many generations. Default is 10. Use None to run all generations.
min_delta (float) – float. Minimum fitness improvement required to reset patience. Default is 1e-4.
- Returns:
the fitted classifier.
- p_value_validation(bootstrap_size=100)[source]#
Computes the permutation and bootstrapping p-values for the classifier and its rules.
- Parameters:
bootstrap_size (int) – integer. Number of bootstraps samples to use.
- load_master_rule_base(rule_base)[source]#
Loads a master rule base to be used in the prediction process.
- Parameters:
rule_base (MasterRuleBase) – ruleBase object.
- Returns:
None
- Return type:
None
- explainable_predict(X, out_class_names=False)[source]#
Returns the predicted class for each sample, with the winning rule, its association degree and confidence interval.
- Parameters:
X (array) – np array samples x features.
out_class_names – if True, the predictions are the consequent names instead of the fitted labels.
- Returns:
predictions, winning rules, winning association degrees and confidence intervals.
- Return type:
an ExplainedPrediction named tuple
- forward(X, out_class_names=False)[source]#
Returns the predicted class for each sample.
- Parameters:
X (array) – np array samples x features.
out_class_names – if True, the output will be the consequent names instead of the fitted labels.
- Returns:
np array samples (x 1) with the predicted class.
- Return type:
array
- predict(X, out_class_names=False)[source]#
Returns the predicted class for each sample.
A fitted classifier predicts the labels it was fitted with (see
classes_); samples where no rule fires are -1 for numeric labels and ‘Unknown’ otherwise. A classifier built from precomputed rules predicts consequent indexes.- Parameters:
X (array) – np array samples x features.
out_class_names – if True, the output will be the consequent names instead of the fitted labels.
- Returns:
np array samples (x 1) with the predicted class.
- Return type:
array
- predict_proba_rules(X, truth_degrees=True)[source]#
Returns the predicted class probabilities for each sample.
- Parameters:
X (array) – np array samples x features.
truth_degrees (bool) – if True, the output will be the truth degrees of the rules. If false, will return the association degrees i.e. the truth degree multiplied by the weights/dominance of the rules. (depending on the inference mode chosen)
- Returns:
np array samples x classes with the predicted class probabilities.
- Return type:
array
- predict_membership_class(X)[source]#
Returns the predicted class memberships for each sample.
- Parameters:
X (array) – np array samples x features.
- Returns:
np array samples x classes with the predicted class probabilities.
- Return type:
array
- predict_proba(X)[source]#
Returns the predicted class probabilities for each sample.
- Parameters:
X (array) – np array samples x features.
- Returns:
np array samples x classes with the predicted class probabilities.
- Return type:
array
- print_rules(return_rules=False, bootstrap_results=False)[source]#
Print the rules contained in the fitted rulebase.
- rename_fuzzy_variables()[source]#
Renames the linguist labels so that high, low and so on are consistent. It does so usually after an optimization process.
- Returns:
None. Names are sorted accorded to the central point of the fuzzy memberships.
- Return type:
None
- get_rulebase()[source]#
Get the rulebase obtained after fitting the classifier to the data.
- Returns:
a matrix format for the rulebase.
- Return type:
list[array]
- reparametrize_loss(alpha, beta)[source]#
Changes the parameters in the loss function.
- Parameters:
alpha (float) – controls the MCC term.
beta (float) – controls the average rule size loss.
Note
Does not check for convexity preservation. The user can play with these parameters as it wills.
- __call__(X)[source]#
Returns the predicted class for each sample.
- Parameters:
X (array) – np array samples x features.
- Returns:
np array samples (x 1) with the predicted class.
- Return type:
array
- set_fit_request(*, bootstrap_size='$UNCHANGED$', candidate_rules='$UNCHANGED$', checkpoint_callback='$UNCHANGED$', checkpoint_path='$UNCHANGED$', checkpoints='$UNCHANGED$', initial_rules='$UNCHANGED$', min_delta='$UNCHANGED$', mutation_eta='$UNCHANGED$', n_gen='$UNCHANGED$', p_value_compute='$UNCHANGED$', patience='$UNCHANGED$', pop_size='$UNCHANGED$', random_state='$UNCHANGED$', sbx_eta='$UNCHANGED$', tournament_size='$UNCHANGED$', var_prob='$UNCHANGED$')#
Configure whether metadata should be requested to be passed to the
fitmethod.Note that this method is only relevant when this estimator is used as a sub-estimator within a meta-estimator and metadata routing is enabled with
enable_metadata_routing=True(seesklearn.set_config()). Please check the User Guide on how the routing mechanism works.The options for each parameter are:
True: metadata is requested, and passed tofitif provided. The request is ignored if metadata is not provided.False: metadata is not requested and the meta-estimator will not pass it tofit.None: metadata is not requested, and the meta-estimator will raise an error if the user provides it.str: metadata should be passed to the meta-estimator with this given alias instead of the original name.
The default (
sklearn.utils.metadata_routing.UNCHANGED) retains the existing request. This allows you to change the request for some parameters and not others.Added in version 1.3.
- Parameters:
bootstrap_size (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for
bootstrap_sizeparameter infit.candidate_rules (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for
candidate_rulesparameter infit.checkpoint_callback (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for
checkpoint_callbackparameter infit.checkpoint_path (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for
checkpoint_pathparameter infit.checkpoints (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for
checkpointsparameter infit.initial_rules (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for
initial_rulesparameter infit.min_delta (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for
min_deltaparameter infit.mutation_eta (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for
mutation_etaparameter infit.n_gen (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for
n_genparameter infit.p_value_compute (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for
p_value_computeparameter infit.patience (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for
patienceparameter infit.pop_size (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for
pop_sizeparameter infit.random_state (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for
random_stateparameter infit.sbx_eta (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for
sbx_etaparameter infit.tournament_size (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for
tournament_sizeparameter infit.var_prob (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for
var_probparameter infit.
- Returns:
self – The updated object.
- Return type:
object
- set_predict_request(*, out_class_names='$UNCHANGED$')#
Configure whether metadata should be requested to be passed to the
predictmethod.Note that this method is only relevant when this estimator is used as a sub-estimator within a meta-estimator and metadata routing is enabled with
enable_metadata_routing=True(seesklearn.set_config()). Please check the User Guide on how the routing mechanism works.The options for each parameter are:
True: metadata is requested, and passed topredictif provided. The request is ignored if metadata is not provided.False: metadata is not requested and the meta-estimator will not pass it topredict.None: metadata is not requested, and the meta-estimator will raise an error if the user provides it.str: metadata should be passed to the meta-estimator with this given alias instead of the original name.
The default (
sklearn.utils.metadata_routing.UNCHANGED) retains the existing request. This allows you to change the request for some parameters and not others.Added in version 1.3.
- Parameters:
out_class_names (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for
out_class_namesparameter inpredict.- Returns:
self – The updated object.
- Return type:
object
- set_score_request(*, sample_weight='$UNCHANGED$')#
Configure whether metadata should be requested to be passed to the
scoremethod.Note that this method is only relevant when this estimator is used as a sub-estimator within a meta-estimator and metadata routing is enabled with
enable_metadata_routing=True(seesklearn.set_config()). Please check the User Guide on how the routing mechanism works.The options for each parameter are:
True: metadata is requested, and passed toscoreif provided. The request is ignored if metadata is not provided.False: metadata is not requested and the meta-estimator will not pass it toscore.None: metadata is not requested, and the meta-estimator will raise an error if the user provides it.str: metadata should be passed to the meta-estimator with this given alias instead of the original name.
The default (
sklearn.utils.metadata_routing.UNCHANGED) retains the existing request. This allows you to change the request for some parameters and not others.Added in version 1.3.
- Parameters:
sample_weight (str, True, False, or None, default=sklearn.utils.metadata_routing.UNCHANGED) – Metadata routing for
sample_weightparameter inscore.- Returns:
self – The updated object.
- Return type:
object
FitRuleBase#
- class ex_fuzzy.evolutionary_fit.FitRuleBase(X, y, nRules, nAnts, n_classes, thread_runner=None, linguistic_variables=None, n_linguistic_variables=3, fuzzy_type=FUZZY_SETS.t1, domain=None, categorical_mask=None, tolerance=0.01, alpha=0.0, beta=0.0, ds_mode=0, allow_unknown=False, backend_name='pymoo', var_names=None)[source]#
Bases:
ProblemClass to model, independently of the optimizer, the fitting of a rulebase for a classification problem using Evolutionary strategies. Supports type 1 and iv fs (iv-type 2)
- vl_names = [[], [], ['Low', 'High'], ['Low', 'Medium', 'High'], ['Low', 'Medium', 'High', 'Very High'], ['Very Low', 'Low', 'Medium', 'High', 'Very High']]#
- __init__(X, y, nRules, nAnts, n_classes, thread_runner=None, linguistic_variables=None, n_linguistic_variables=3, fuzzy_type=FUZZY_SETS.t1, domain=None, categorical_mask=None, tolerance=0.01, alpha=0.0, beta=0.0, ds_mode=0, allow_unknown=False, backend_name='pymoo', var_names=None)[source]#
Cosntructor method. Initializes the classifier with the number of antecedents, linguist variables and the kind of fuzzy set desired.
- Parameters:
X (array) – np array or pandas dataframe samples x features.
y (array) – np vector containing the target classes. vector sample
nRules (int) – number of rules to optimize.
nAnts (int) – max number of antecedents to use.
n_class – number of classes in the problem. If None (as default) it will be computed from the data.
linguistic_variables (list[fuzzyVariable]) – list of linguistic variables precomputed. If given, the rest of conflicting arguments are ignored.
n_linguistic_variables (int) – number of linguistic variables per antecedent.
fuzzy_type – Define the fuzzy set or fuzzy set extension used as linguistic variable.
domain (list) – list with the upper and lower domains of each input variable. If None (as default) it will stablish the empirical min/max as the limits.
tolerance (float) – float. Tolerance for the size evaluation.
alpha (float) – float. Weight for the rulebase size term in the fitness function. (Penalizes number of rules)
beta (float) – float. Weight for the average rule size term in the fitness function.
ds_mode (int) – int. Mode for the dominance score. 0: normal dominance score, 1: rules without weights, 2: weights optimized for each rule based on the data.
allow_unknown (bool) – if True, the classifier will allow the unknown class in the classification process. (Which would be a -1 value)
var_names (list) – list of variable names. If None, extracted from DataFrame columns or auto-generated.
- encode_rulebase(rule_base, optimize_lv)[source]#
Given a rule base, constructs the corresponding gene associated with that rule base.
GENE STRUCTURE
First: antecedents chosen by each rule. Size: nAnts * nRules (index of the antecedent) Second: Variable linguistics used. Size: nAnts * nRules Third: Parameters for the fuzzy partitions of the chosen variables. Size: nAnts * self.n_linguistic_variables * 8|4 (2 trapezoidal memberships if t2) Four: Consequent classes. Size: nRules
- Parameters:
rule_base (MasterRuleBase) – rule base object.
optimize_lv (bool) – must be False: only genes over fixed linguistic variables can be encoded.
- Returns:
np array of size self.single_gen_size.
- Raises:
NotImplementedError – if optimize_lv is True.
ValueError – if the problem does not have one antecedent slot per feature.
- Return type:
array
- array_evaluation = True#
Set to False to force the object decoder, for benchmarks and parity tests.
- torch_devices = ('cuda',)#
Device types on which populations may be scored by the exact PyTorch objective. Scoring CPU tensors gains nothing over the NumPy routes, so only CUDA is enabled; parity tests add ‘cpu’ to run the same code.
- fitness_func(ruleBase, X, y, tolerance, alpha=0.0, beta=0.0, precomputed_truth=None)[source]#
Fitness function for the optimization problem.
- Parameters:
ruleBase (RuleBase) – RuleBase object
X (array) – array of train samples. X shape = (n_samples, n_features)
y (array) – array of train labels. y shape = (n_samples,)
tolerance (float) – float. Tolerance for the size evaluation.
alpha (float) – float. Weight for the accuracy term.
beta (float) – float. Weight for the average rule size term.
precomputed_truth (array) – np array. If given, it will be used as the truth values for the evaluation.
- Returns:
float. Fitness value.
- Return type:
float
See Also#
ex_fuzzy.classifiersex_fuzzy.rulesex_fuzzy.rule_miningex_fuzzy.eval_tools