API Reference

SnapBoostClassifier / SnapBoostRegressor

Recommended entry points (analogous to XGBClassifier / XGBRegressor).

from snapboost import SnapBoostClassifier, SnapBoostRegressor

clf = SnapBoostClassifier(
    num_iterations=100,
    learning_rate=0.1,
    p_tree=0.9,
    min_max_depth=2,
    max_max_depth=4,
    alpha=1.0,
    gamma=1.0,
    random_state=42,
    verbose=True,
)
clf.fit(X, y)

reg = SnapBoostRegressor(num_iterations=100, random_state=42)
reg.fit(X, y)

Methods

Method

Classifier

Regressor

Description

fit(X, y, sample_weight=None, eval_set=None, *, eval_sample_weight=None)

✓

✓

Train, optionally with weights and validation data

predict(X)

✓

✓

Class labels from classes_ or continuous values

predict_proba(X)

✓

Class probabilities, shape (n_samples, n_classes)

decision_function(X)

✓

Raw logits: (n_samples,) binary, (n_samples, n_classes) multiclass

staged_predict(X)

✓

✓

Predictions after each boosting round

permutation_importance(X, y)

✓

✓

Permutation importance of original features

score(X, y)

✓

✓

Accuracy or R²

evaluate(X, y)

✓

✓

Prints and returns log loss or RMSE

class snapboost.SnapBoostClassifier(num_iterations=100, learning_rate=0.1, p_tree=0.9, p_linear=0.0, min_max_depth=2, max_max_depth=4, min_samples_leaf=10, alpha=1.0, alpha_linear=None, gamma=1.0, n_components=100, max_features=None, scale_features=True, kernel_gammas=None, kernel_types=('rbf',), monotonic_cst=None, random_state=None, verbose=False, selection_strategy='random', line_search=False, early_stopping_rounds=None, min_delta=0.0, subsample=1.0, objective='auto', objective_parameter=None)[source]

Bases: _SnapBoostMixin, HNBMClassifier

SnapBoost for classification.

A heterogeneous Newton boosting machine that uses decision trees and random Fourier feature ridge regressors. Binary targets use logistic loss. Multiclass targets use softmax Newton boosting with one scalar learner per class each round. Supports random or greedy learner selection, per-round line search, subsampling, observation weights, validation histories, and early stopping.

set_params(**params)[source]

Set the parameters of this estimator.

The method works on simple estimators as well as on nested objects (such as Pipeline). The latter have parameters of the form <component>__<parameter> so that it’s possible to update each component of a nested object.

Parameters:

**params (dict) – Estimator parameters.

Returns:

self – Estimator instance.

Return type:

estimator instance

decision_function(X)

Return classification logits.

Binary models return shape (n_samples,). Multiclass models return shape (n_samples, n_classes).

evaluate(X, y)

Print and return log loss (classification) or RMSE (regression).

fit(X, y, sample_weight=None, eval_set=None, eval_metric=None, callbacks=None, candidate_n_jobs=1, *, eval_sample_weight=None)

Train the model.

Parameters:
  • X (array-like of shape (n_samples, n_features)) – Feature matrix.

  • y (array-like of shape (n_samples,)) – Target values.

  • sample_weight (array-like of shape (n_samples,), default=None) – Non-negative observation weights.

  • eval_set (tuple (X_validation, y_validation), default=None) – Optional validation pair used for history and early stopping.

  • eval_metric (callable, default=None) – Optional metric(y, raw_prediction) -> float recorded per round. y is the original label vector, not the internal Newton encoding.

  • callbacks (iterable of callable, default=None) – Functions called with a per-round state dictionary. Returning a truthy value requests an orderly stop after the current round. The round that triggers early stopping is reported too, before the ensemble is rolled back to best_iteration_.

  • candidate_n_jobs (int, default=1) – Threads used to fit greedy candidates. Has no effect for random selection. -1 uses all available logical CPUs.

  • eval_sample_weight (array-like of shape (n_eval,), default=None) – Non-negative weights for eval_set. Requires eval_set.

Return type:

self

permutation_importance(X, y, *, n_repeats=5, random_state=None, scoring=None, n_jobs=None, sample_weight=None, max_samples=1.0)

Permutation importance of the original features.

This is the right importance measure for mixed tree and smooth learners. It returns the sklearn.inspection.permutation_importance() result.

predict(X)

Predict using the model.

Classification returns labels from classes_; regression returns continuous values.

predict_proba(X)

Predict class probabilities (classification mode only).

Returns:

Probabilities in classes_ order.

Return type:

ndarray of shape (n_samples, n_classes)

score(X, y, sample_weight=None)

Return accuracy on provided data and labels.

In multi-label classification, this is the subset accuracy which is a harsh metric since you require for each sample that each label set be correctly predicted.

Parameters:
  • X (array-like of shape (n_samples, n_features)) – Test samples.

  • y (array-like of shape (n_samples,) or (n_samples, n_outputs)) – True labels for X.

  • sample_weight (array-like of shape (n_samples,), default=None) – Sample weights.

Returns:

score – Mean accuracy of self.predict(X) w.r.t. y.

Return type:

float

staged_decision_function(X)

Yield classification logits after each boosting round.

staged_predict(X)

Yield predictions after each boosting round.

The first value includes the first fitted learner. The last value matches predict().

staged_predict_proba(X)

Yield class probabilities after each boosting round.

class snapboost.SnapBoostRegressor(num_iterations=100, learning_rate=0.1, p_tree=0.9, p_linear=0.0, min_max_depth=2, max_max_depth=4, min_samples_leaf=10, alpha=1.0, alpha_linear=None, gamma=1.0, n_components=100, max_features=None, scale_features=True, kernel_gammas=None, kernel_types=('rbf',), monotonic_cst=None, random_state=None, verbose=False, selection_strategy='random', line_search=False, early_stopping_rounds=None, min_delta=0.0, subsample=1.0, objective='auto', objective_parameter=None)[source]

Bases: _SnapBoostMixin, HNBMRegressor

SnapBoost for regression.

A heterogeneous Newton boosting machine that uses decision trees and random Fourier feature ridge regressors. Supports random or greedy learner selection, per-round line search, subsampling, observation weights, validation histories, and early stopping.

set_params(**params)[source]

Set the parameters of this estimator.

The method works on simple estimators as well as on nested objects (such as Pipeline). The latter have parameters of the form <component>__<parameter> so that it’s possible to update each component of a nested object.

Parameters:

**params (dict) – Estimator parameters.

Returns:

self – Estimator instance.

Return type:

estimator instance

evaluate(X, y)

Print and return log loss (classification) or RMSE (regression).

fit(X, y, sample_weight=None, eval_set=None, eval_metric=None, callbacks=None, candidate_n_jobs=1, *, eval_sample_weight=None)

Train the model.

Parameters:
  • X (array-like of shape (n_samples, n_features)) – Feature matrix.

  • y (array-like of shape (n_samples,)) – Target values.

  • sample_weight (array-like of shape (n_samples,), default=None) – Non-negative observation weights.

  • eval_set (tuple (X_validation, y_validation), default=None) – Optional validation pair used for history and early stopping.

  • eval_metric (callable, default=None) – Optional metric(y, raw_prediction) -> float recorded per round. y is the original label vector, not the internal Newton encoding.

  • callbacks (iterable of callable, default=None) – Functions called with a per-round state dictionary. Returning a truthy value requests an orderly stop after the current round. The round that triggers early stopping is reported too, before the ensemble is rolled back to best_iteration_.

  • candidate_n_jobs (int, default=1) – Threads used to fit greedy candidates. Has no effect for random selection. -1 uses all available logical CPUs.

  • eval_sample_weight (array-like of shape (n_eval,), default=None) – Non-negative weights for eval_set. Requires eval_set.

Return type:

self

permutation_importance(X, y, *, n_repeats=5, random_state=None, scoring=None, n_jobs=None, sample_weight=None, max_samples=1.0)

Permutation importance of the original features.

This is the right importance measure for mixed tree and smooth learners. It returns the sklearn.inspection.permutation_importance() result.

predict(X)

Predict using the model.

Classification returns labels from classes_; regression returns continuous values.

score(X, y, sample_weight=None)

Return coefficient of determination on test data.

The coefficient of determination, \(R^2\), is defined as \((1 - \frac{u}{v})\), where \(u\) is the residual sum of squares ((y_true - y_pred)** 2).sum() and \(v\) is the total sum of squares ((y_true - y_true.mean()) ** 2).sum(). The best possible score is 1.0 and it can be negative (because the model can be arbitrarily worse). A constant model that always predicts the expected value of y, disregarding the input features, would get a \(R^2\) score of 0.0.

Parameters:
  • X (array-like of shape (n_samples, n_features)) – Test samples. For some estimators this may be a precomputed kernel matrix or a list of generic objects instead with shape (n_samples, n_samples_fitted), where n_samples_fitted is the number of samples used in the fitting for the estimator.

  • y (array-like of shape (n_samples,) or (n_samples, n_outputs)) – True values for X.

  • sample_weight (array-like of shape (n_samples,), default=None) – Sample weights.

Returns:

score – \(R^2\) of self.predict(X) w.r.t. y.

Return type:

float

Notes

The \(R^2\) score used when calling score on a regressor uses multioutput='uniform_average' from version 0.23 to keep consistent with default value of r2_score(). This influences the score method of all the multioutput regressors (except for MultiOutputRegressor).

staged_predict(X)

Yield predictions after each boosting round.

The first value includes the first fitted learner. The last value matches predict().

Fitted adaptive models expose base_score_, learner_weights_, history_, best_iteration_, n_iter_, and for classifiers classes_ / n_classes_. Multiclass ensemble_ entries are one fitted scalar learner per class. When early_stopping_rounds triggers, the best validation ensemble is restored before fit returns. Passing an eval_set without early_stopping_rounds still records best_iteration_, but no learners are discarded, so predictions use all n_iter_ of them.

Models also inherit compact(min_abs_weight=0.0, inplace=False) from HNBM. Compaction is explicit and never runs automatically.

Tabular preprocessing

make_tabular_preprocessor() returns a dense scikit-learn ColumnTransformer that median-imputes numeric data, adds optional missingness indicators, and imputes/one-hot encodes categorical data with unseen-category support. Use it in a normal Pipeline; it does not modify SnapBoost internals.

SnapBoost (deprecated)

Accepts a mode parameter ("classification" or "regression"). This class emits FutureWarning and will be removed in 2.0. Prefer the task-specific classes above.

from snapboost import SnapBoost

model = SnapBoost(
    num_iterations=100,
    learning_rate=0.1,
    p_tree=0.9,
    mode="classification",
    random_state=42,
)
model.fit(X, y)
class snapboost.SnapBoost(num_iterations=100, learning_rate=0.1, p_tree=0.9, p_linear=0.0, min_max_depth=2, max_max_depth=4, min_samples_leaf=10, alpha=1.0, alpha_linear=None, gamma=1.0, n_components=100, mode='classification', random_state=None, verbose=False)[source]

Bases: _SnapBoostMixin, HNBM

HNBM realization using decision trees and RFF ridge regressors.

Deprecated. Prefer SnapBoostClassifier or SnapBoostRegressor for task-specific models without a mode parameter.

set_params(**params)[source]

Set the parameters of this estimator.

The method works on simple estimators as well as on nested objects (such as Pipeline). The latter have parameters of the form <component>__<parameter> so that it’s possible to update each component of a nested object.

Parameters:

**params (dict) – Estimator parameters.

Returns:

self – Estimator instance.

Return type:

estimator instance

decision_function(X)

Return classification logits.

Binary models return shape (n_samples,). Multiclass models return shape (n_samples, n_classes).

evaluate(X, y)

Print and return log loss (classification) or RMSE (regression).

fit(X, y, sample_weight=None, eval_set=None, eval_metric=None, callbacks=None, candidate_n_jobs=1, *, eval_sample_weight=None)

Train the model.

Parameters:
  • X (array-like of shape (n_samples, n_features)) – Feature matrix.

  • y (array-like of shape (n_samples,)) – Target values.

  • sample_weight (array-like of shape (n_samples,), default=None) – Non-negative observation weights.

  • eval_set (tuple (X_validation, y_validation), default=None) – Optional validation pair used for history and early stopping.

  • eval_metric (callable, default=None) – Optional metric(y, raw_prediction) -> float recorded per round. y is the original label vector, not the internal Newton encoding.

  • callbacks (iterable of callable, default=None) – Functions called with a per-round state dictionary. Returning a truthy value requests an orderly stop after the current round. The round that triggers early stopping is reported too, before the ensemble is rolled back to best_iteration_.

  • candidate_n_jobs (int, default=1) – Threads used to fit greedy candidates. Has no effect for random selection. -1 uses all available logical CPUs.

  • eval_sample_weight (array-like of shape (n_eval,), default=None) – Non-negative weights for eval_set. Requires eval_set.

Return type:

self

predict(X)

Predict using the model.

Classification returns labels from classes_; regression returns continuous values.

predict_proba(X)

Predict class probabilities (classification mode only).

Returns:

Probabilities in classes_ order.

Return type:

ndarray of shape (n_samples, n_classes)

score(X, y)

Return accuracy (classification) or R² (regression).

RandomFourierRidgeRegressor

Ridge regression on random Fourier features approximating RBF or Laplacian kernels. Used as a non-tree learner in the SnapBoost pool.

class snapboost.RandomFourierRidgeRegressor(alpha=1.0, gamma=1.0, n_components=100, random_state=None, scale_features=True, kernel='rbf')[source]

Bases: BaseEstimator, RegressorMixin

Ridge regression on random Fourier features approximating an RBF kernel.

Matches the linear + RFF base learner used in the original SnapBoost paper, scaling linearly in the number of samples instead of exact KernelRidge. Features are standardized by default because RBF distances are sensitive to input scale. Set scale_features=False for pre-scaled inputs.

fit(X, y, sample_weight=None)[source]
predict(X)[source]

Exact kernel ridge estimators

SnapBoostKernelRidgeClassifier and SnapBoostKernelRidgeRegressor swap the random Fourier feature learner for an exact RBF kernel ridge learner. They keep a frozen constructor surface (num_iterations, learning_rate, p_tree, min_max_depth, max_max_depth, min_samples_leaf, alpha, gamma, random_state, verbose) and so do not expose greedy selection, line search, subsampling, or early stopping. Classification inherits binary logistic loss and multiclass softmax from HNBM. Memory grows quadratically in the number of samples, so prefer the RFF path unless an exact kernel is required.

from snapboost import SnapBoostKernelRidgeClassifier

clf = SnapBoostKernelRidgeClassifier(num_iterations=50, gamma=0.5, random_state=42)
clf.fit(X, y)
class snapboost.SnapBoostKernelRidgeClassifier(num_iterations=100, learning_rate=0.1, p_tree=0.9, min_max_depth=2, max_max_depth=4, min_samples_leaf=10, alpha=1.0, gamma=1.0, random_state=None, verbose=False)[source]

Bases: _KernelRidgePoolMixin, HNBMClassifier

SnapBoost classifier using exact RBF kernel ridge learners.

Binary targets use logistic loss; multiclass targets use softmax Newton boosting. This specialized surface is frozen in 1.0: it does not expose the adaptive training controls of SnapBoostClassifier. Prefer the RFF-based classifier unless an exact kernel is required.

set_params(**params)[source]

Set the parameters of this estimator.

The method works on simple estimators as well as on nested objects (such as Pipeline). The latter have parameters of the form <component>__<parameter> so that it’s possible to update each component of a nested object.

Parameters:

**params (dict) – Estimator parameters.

Returns:

self – Estimator instance.

Return type:

estimator instance

decision_function(X)

Return classification logits.

Binary models return shape (n_samples,). Multiclass models return shape (n_samples, n_classes).

evaluate(X, y)

Print and return log loss (classification) or RMSE (regression).

fit(X, y, sample_weight=None, eval_set=None, eval_metric=None, callbacks=None, candidate_n_jobs=1, *, eval_sample_weight=None)

Train the model.

Parameters:
  • X (array-like of shape (n_samples, n_features)) – Feature matrix.

  • y (array-like of shape (n_samples,)) – Target values.

  • sample_weight (array-like of shape (n_samples,), default=None) – Non-negative observation weights.

  • eval_set (tuple (X_validation, y_validation), default=None) – Optional validation pair used for history and early stopping.

  • eval_metric (callable, default=None) – Optional metric(y, raw_prediction) -> float recorded per round. y is the original label vector, not the internal Newton encoding.

  • callbacks (iterable of callable, default=None) – Functions called with a per-round state dictionary. Returning a truthy value requests an orderly stop after the current round. The round that triggers early stopping is reported too, before the ensemble is rolled back to best_iteration_.

  • candidate_n_jobs (int, default=1) – Threads used to fit greedy candidates. Has no effect for random selection. -1 uses all available logical CPUs.

  • eval_sample_weight (array-like of shape (n_eval,), default=None) – Non-negative weights for eval_set. Requires eval_set.

Return type:

self

permutation_importance(X, y, *, n_repeats=5, random_state=None, scoring=None, n_jobs=None, sample_weight=None, max_samples=1.0)

Permutation importance of the original features.

This is the right importance measure for mixed tree and smooth learners. It returns the sklearn.inspection.permutation_importance() result.

predict(X)

Predict using the model.

Classification returns labels from classes_; regression returns continuous values.

predict_proba(X)

Predict class probabilities (classification mode only).

Returns:

Probabilities in classes_ order.

Return type:

ndarray of shape (n_samples, n_classes)

score(X, y, sample_weight=None)

Return accuracy on provided data and labels.

In multi-label classification, this is the subset accuracy which is a harsh metric since you require for each sample that each label set be correctly predicted.

Parameters:
  • X (array-like of shape (n_samples, n_features)) – Test samples.

  • y (array-like of shape (n_samples,) or (n_samples, n_outputs)) – True labels for X.

  • sample_weight (array-like of shape (n_samples,), default=None) – Sample weights.

Returns:

score – Mean accuracy of self.predict(X) w.r.t. y.

Return type:

float

staged_decision_function(X)

Yield classification logits after each boosting round.

staged_predict(X)

Yield predictions after each boosting round.

The first value includes the first fitted learner. The last value matches predict().

staged_predict_proba(X)

Yield class probabilities after each boosting round.

class snapboost.SnapBoostKernelRidgeRegressor(num_iterations=100, learning_rate=0.1, p_tree=0.9, min_max_depth=2, max_max_depth=4, min_samples_leaf=10, alpha=1.0, gamma=1.0, random_state=None, verbose=False)[source]

Bases: _KernelRidgePoolMixin, HNBMRegressor

SnapBoost regressor using exact RBF kernel ridge learners.

This specialized surface is frozen in 1.0: it does not expose the adaptive training controls of SnapBoostRegressor. Prefer the RFF-based regressor unless an exact kernel is required.

set_params(**params)[source]

Set the parameters of this estimator.

The method works on simple estimators as well as on nested objects (such as Pipeline). The latter have parameters of the form <component>__<parameter> so that it’s possible to update each component of a nested object.

Parameters:

**params (dict) – Estimator parameters.

Returns:

self – Estimator instance.

Return type:

estimator instance

evaluate(X, y)

Print and return log loss (classification) or RMSE (regression).

fit(X, y, sample_weight=None, eval_set=None, eval_metric=None, callbacks=None, candidate_n_jobs=1, *, eval_sample_weight=None)

Train the model.

Parameters:
  • X (array-like of shape (n_samples, n_features)) – Feature matrix.

  • y (array-like of shape (n_samples,)) – Target values.

  • sample_weight (array-like of shape (n_samples,), default=None) – Non-negative observation weights.

  • eval_set (tuple (X_validation, y_validation), default=None) – Optional validation pair used for history and early stopping.

  • eval_metric (callable, default=None) – Optional metric(y, raw_prediction) -> float recorded per round. y is the original label vector, not the internal Newton encoding.

  • callbacks (iterable of callable, default=None) – Functions called with a per-round state dictionary. Returning a truthy value requests an orderly stop after the current round. The round that triggers early stopping is reported too, before the ensemble is rolled back to best_iteration_.

  • candidate_n_jobs (int, default=1) – Threads used to fit greedy candidates. Has no effect for random selection. -1 uses all available logical CPUs.

  • eval_sample_weight (array-like of shape (n_eval,), default=None) – Non-negative weights for eval_set. Requires eval_set.

Return type:

self

permutation_importance(X, y, *, n_repeats=5, random_state=None, scoring=None, n_jobs=None, sample_weight=None, max_samples=1.0)

Permutation importance of the original features.

This is the right importance measure for mixed tree and smooth learners. It returns the sklearn.inspection.permutation_importance() result.

predict(X)

Predict using the model.

Classification returns labels from classes_; regression returns continuous values.

score(X, y, sample_weight=None)

Return coefficient of determination on test data.

The coefficient of determination, \(R^2\), is defined as \((1 - \frac{u}{v})\), where \(u\) is the residual sum of squares ((y_true - y_pred)** 2).sum() and \(v\) is the total sum of squares ((y_true - y_true.mean()) ** 2).sum(). The best possible score is 1.0 and it can be negative (because the model can be arbitrarily worse). A constant model that always predicts the expected value of y, disregarding the input features, would get a \(R^2\) score of 0.0.

Parameters:
  • X (array-like of shape (n_samples, n_features)) – Test samples. For some estimators this may be a precomputed kernel matrix or a list of generic objects instead with shape (n_samples, n_samples_fitted), where n_samples_fitted is the number of samples used in the fitting for the estimator.

  • y (array-like of shape (n_samples,) or (n_samples, n_outputs)) – True values for X.

  • sample_weight (array-like of shape (n_samples,), default=None) – Sample weights.

Returns:

score – \(R^2\) of self.predict(X) w.r.t. y.

Return type:

float

Notes

The \(R^2\) score used when calling score on a regressor uses multioutput='uniform_average' from version 0.23 to keep consistent with default value of r2_score(). This influences the score method of all the multioutput regressors (except for MultiOutputRegressor).

staged_predict(X)

Yield predictions after each boosting round.

The first value includes the first fitted learner. The last value matches predict().

SnapBoost_KernelRidge is the deprecated mode-based equivalent. It emits FutureWarning and will be removed in 2.0.

Optional learner and preprocessing helpers

class snapboost.WeightedLinearRegressor(alpha=1.0, scale_features=True)[source]

Bases: BaseEstimator, RegressorMixin

Standardized weighted ridge regression for Newton working targets.

class snapboost.WeightedKernelRidgeRegressor(alpha=1.0, gamma=1.0, scale_features=True)[source]

Bases: BaseEstimator, RegressorMixin

Standardized weighted RBF kernel ridge for Newton working targets.

class snapboost.LaplacianSampler(gamma=1.0, n_components=100, random_state=None)[source]

Bases: BaseEstimator, TransformerMixin

Random Fourier features for the Laplacian (L1 exponential) kernel.

snapboost.make_tabular_preprocessor(categorical_features=(), *, scale_numeric=False, add_missing_indicators=True)[source]

Build a dense, leakage-safe preprocessor for use before SnapBoost.

Numeric columns are median-imputed, optionally augmented with missingness indicators and standardized. Categorical columns are most-frequent-imputed and one-hot encoded with unknown-category handling. The returned transformer belongs in a normal scikit-learn Pipeline and does not alter SnapBoost’s internal training or model format.

categorical_features may be a single column name or index, or a sequence of names/indices. A string is treated as one column name, not as a sequence of characters.

Parameters:
  • categorical_features (object)

  • scale_numeric (bool)

  • add_missing_indicators (bool)

Return type:

ColumnTransformer

HNBM

Abstract base classes for custom heterogeneous ensembles are provided by the hnbm package. Subclass and configure base_learners_ / probabilities_ before calling fit:

from sklearn.tree import DecisionTreeRegressor
from hnbm import HNBMClassifier

class MyClassifier(HNBMClassifier):
    def __init__(self, **kwargs):
        super().__init__(**kwargs)
        self.base_learners_ = [DecisionTreeRegressor(max_depth=5)]
        self.probabilities_ = [1.0]