API Reference#
Created on Thu Jun 6 13:28:15 2024
@author: Polina Arsenteva
- kaov.kaov.ordered_eigsy(matrix, eps=None, clip=True)[source]#
Calculates the eigendecomposition of the matrix, using a solver implemented in C++.
Eigen values and eigen vectors are stored by decreasing eigen value order.
Note 1: only positive non-null eigen values and corresponding eigen vectors are returned. Eigen values lower than eps threshold a clipped to zeros.
Note 2: input matrix is assumed to be symmetrical and positive semi-definite, such that its eigen values are positive or null. It may happen due to numerical issues that the eigen decomposition finds negative eigen values for such matrix anyway. These are also clipped to 0.
- Parameters:
- matrix2-d array_like
Matrix to decompose.
- epsfloat, optional
minimum threshold value to clip lower eigen values to zeros. If None (default), then machine precision (given by torch.finfo()) for matrix dtype is used as threshold.
- clipboolean,
flag to enable/disable eigen value clipping.
- Returns:
- sptorch.Tensor
Eigenvalues.
- evtorch.Tensor
Eigenvectors.
- kaov.kaov.convert_to_torch(A)[source]#
Converts A to torch.Tensor.
- Parameters:
- Aarray_like
Container to convert.
- Returns:
- Btorch.Tensor
A converted to torch.
- kaov.kaov.distances(x, y=None)[source]#
Computes the distances between each pair of the two collections of row vectors of x and y, or x and itself if y is not provided.
- Parameters:
- xtorch.Tensor
Input 2-d tensor.
- yNone or torch.Tensor
Input 2-d tensor. Replaced by x if None.
- Returns:
- sq_diststorch.Tensor
Distance matrix.
- kaov.kaov.linear_kernel(x, y=None)[source]#
Computes the standard linear kernel k(x,y)= <x,y>.
- Parameters:
- xtorch.Tensor
2-d tensor containing the data to kernalize.
- yNone or torch.Tensor
2-d tensor containing the data to kernalize. Replaced by x if None.
- Returns:
- Ktorch.Tensor
Kernel matrix (gram).
- kaov.kaov.gauss_kernel(x, y=None, sigma=1)[source]#
Computes the standard Gaussian kernel k(x,y)=exp(- ||x-y||**2 / (2 * sigma**2)).
- Parameters:
- xtorch.Tensor
2-d tensor containing the data to kernalize.
- yNone or torch.Tensor
2-d tensor containing the data to kernalize. Replaced by x if None.
- sigmaint
Standard deviation for Gaussian kernel, 1 by default.
- Returns:
- Ktorch.Tensor
Kernel matrix (gram).
- kaov.kaov.gauss_kernel_median(x, y=None, bandwidth='median', median_coef=1, return_bandwidth=False)[source]#
Computes the gaussian kernel with bandwidth set as the median of the distances between pairs of observations (bandwidth=’median’).
- Parameters:
- xtorch.Tensor
2-d tensor containing the data to kernalize.
- yNone or torch.Tensor
2-d tensor containing the data to kernalize. Replaced by x if None.
- bandwidth‘median’ or float
If ‘median’ (default), the bandwidth is calculated with the median method. If float, the value is assigned to the bandwidth.
- median_coeffloat
Coefficient in the badwidth calculation, 1 by default.
- return_bandwidthbool
Bandwidth, calculated with the median method. False by default.
- Returns:
- kernelcallable
Kernel function.
- if return_bandwidth=True:
- computed_bandwidthfloat
Bandwidth, calculated with the median method.
- class kaov.kaov.OneHot(reference=-1)[source]#
Bases:
objectClass defining One Hot encoding: a coding scheme for linear models, along with other well known ones such as Treatment or Difference coding. It is to be integrated in a formula defining the linear model, provided by patsy’s formula interface. It is recommended to use one hot encoding with the testing framework implemented in AOV, especially with more than one factor.
Example of a formula with OneHot:
'y1 + y2 ~ C(x1, OneHot) + C(x2, OneHot)'See Readme and tutorials for kAOV for more details and examples.
- class kaov.kaov.Data(endog, exog, meta=None, endog_names=None, exog_names=None, nystrom=False, n_landmarks=None, random_gen=None)[source]#
Bases:
objectClass containing data-related structures for the kernel analysis of variance.
- Parameters:
- endog2-d array_like
An array_like with dimensions nobs x nvar containing nobs values of nvar dependent variables.
- exog2-d array_like
An array_like with dimensions nobs x nlvl containing nobs values of nlvl independent variables.
- metaNone or 2-d array_like, optional
An array_like with the metadata for the dataset, i.e. containing information on factors. Used for visualizations.
- endog_namesNone or 1-d array_like, optional
A 1-dimensional array_like containing names of the dependent variables. If not specified (default), will be retrieved from endog or assigned to numbers with respect to the order.
- exog_namesNone or 1-d array_like, optional
A 1-dimensional array_like containing names of the independent variables. If not specified (default), will be retrieved from exog or assigned to numbers with respect to the order.
- nystrombool, optional
If True, computes the Nystrom landmarks, in which case the observations in all attributes correspond to the landmarks and not the original data. The default is False.
- n_landmarks: int, optional
Number of landmarks used in the Nystrom method. If unspecified, one fifth of the observations are selected as landmarks.
- random_genint, Generator, RandomState instance or None
Determines random number generation for the landmarks selection. If None (default), the generator is the RandomState instance used by np.random. To ensure the results are reproducible, pass an int to instanciate the seed, or a Generator/RandomState instance (recommended).
- Attributes:
- endog2-d torch.tensor
A tensor with dimensions nobs x nvar containing nobs values of nvar dependent variables.
- exog2-d torch.tensor
A tensor with dimensions nobs x nlvl containing nobs values of nlvl independent variables.
- metaNone or 2-d array_like
An array_like with the metadata for the dataset, i.e. containing information on factors. Used for visualizations. The default is None.
- nobsint
Number of observations. If nystrom=True, corresponds to the number of landmarks n_landmarks.
- nlvlint
Number of independent variables (i.e. levels of all factors).
- nvarint
Number of dependent variables.
- index1-d array_like
Observation labels.
- endog_names1-d array_like
A 1-dimensional array_like containing names of the dependent variables.
- exog_names1-d array_like
A 1-dimensional array_like containing names of the independent variables.
- nystrombool
False by default, True if the Nystrom approximation is performed, in which case the observations in all attributes correspond to the landmarks and not the original data.
- class kaov.kaov.AOV(endog, exog, meta=None, endog_names=None, exog_names=None, nystrom=False, n_landmarks=None, random_gen=None, kernel_function='gauss', kernel_bandwidth='median', kernel_median_coef=1, verbose=1)[source]#
Bases:
objectClass implementing Kernel Analysis Of Variance.
- Parameters:
- endog2-d array_like
An array_like with dimensions nobs x nvar containing nobs values of nvar dependent variables.
- exog2-d array_like
An array_like with dimensions nobs x nlvl containing nobs values of nlvl independent variables.
- metaNone or 2-d array_like, optional
An array_like with the metadata for the dataset, i.e. containing information on factors. Used for visualizations.
- endog_namesNone or 1-d array_like, optional
A 1-dimensional array_like containing names of the dependent variables. If not specified (default), will be retrieved from endog or assigned to numbers with respect to the order.
- exog_namesNone or 1-d array_like, optional
A 1-dimensional array_like containing names of the independent variables. If not specified (default), will be retrieved from exog or assigned to numbers with respect to the order.
- kernel_functioncallable or str, optional
Specifies the kernel function. Acceptable values in the form of a string are ‘gauss’ (default) and ‘linear’. Pass a callable for a user-defined kernel function.
- kernel_bandwidth‘median’ or float, optional
Value of the bandwidth for kernels using a bandwidth. If ‘median’ (default), the bandwidth will be set as the median or its multiple, depending on the value of the parameter median_coef. Pass a float for a user-defined value of the bandwidth.
- kernel_median_coeffloat, optional
Multiple of the median to compute bandwidth if kernel_bandwidth=’median’. The default is 1.
- nystrombool, optional
If True, computes the Nystrom landmarks, in which case the observations in all attributes correspond to the landmarks and not the original data. The default is False.
- n_landmarks: int, optional
Number of landmarks used in the Nystrom method. If unspecified, one fifth of the observations are selected as landmarks.
- random_genint, Generator, RandomState instance or None
Determines random number generation for the landmarks selection. If None (default), the generator is the RandomState instance used by np.random. To ensure the results are reproducible, pass an int to instanciate the seed, or a Generator/RandomState instance (recommended).
- Attributes:
- datainstance of class Data
Contains various information on the original dataset, see the documentation of the class Data for more details.
- data_nystromNone or instance of class Data
Contains various information on the Nystrom dataset, see the documentation of the class Data for more details. If None, Nystrom is not taken into account in the computations.
- kernel_functioncallable or str, optional
Specifies the kernel function.
- kernel_bandwidth‘median’ or float, optional
Value of the bandwidth for kernels using a bandwidth.
- kernel_median_coeffloat, optional
Multiple of the median to compute bandwidth if kernel_bandwidth=’median’. The default is 1.
- kernel: callable
Kernel function used for calculations.
- computed_bandwidthfloat
The value of the kernel bandwidth.
- verboseint, optional
The higher the verbosity, the more messages keeping track of computations. The default is 0. - < 1: no messages, - 1: warnings are printed once, - 2: warnings are printed every time they appear.
Notes
The from_formula interface is the recommended method to specify a model.
- classmethod from_formula(formula, data, kernel_function='gauss', nystrom=False, n_landmarks=None, random_gen=None, verbose=1, kernel_bandwidth='median', kernel_median_coef=1)[source]#
Creates a kernel linear model from a formula and a dataframe.
- Parameters:
- formulastr
The formula specifying the model. For more details on the formula interface see https://patsy.readthedocs.io/en/latest/formulas.html.
- datapandas.DataFrame
The data for the model. Columns must contain the values for the factors in the formula, with the names matching those in the formula.
- kernel_functioncallable or str, optional
Specifies the kernel function. Acceptable values in the form of a string are ‘gauss’ (default) and ‘linear’. Pass a callable for a user-defined kernel function.
- kernel_bandwidth‘median’ or float, optional
Value of the bandwidth for kernels using a bandwidth. If ‘median’ (default), the bandwidth will be set as the median or its multiple, depending on the value of the parameter median_coef. Pass a float for a user-defined value of the bandwidth.
- kernel_median_coeffloat, optional
Multiple of the median to compute bandwidth if kernel_bandwidth=’median’. The default is 1.
- verboseint, optional
The higher the verbosity, the more messages keeping track of computations. The default is 0. - < 1: no messages, - 1: warnings are printed once, - 2: warnings are printed every time they appear.
- Returns:
- aov_objinstance of AOV
Examples
Importing data:
>>> import pandas as pd >>> from kaov import AOV >>> url = "https://raw.githubusercontent.com/LMJL-Alea/kAOV/refs/heads/main/Data/reversion_kAOV.csv" >>> data = pd.read_csv(url, index_col=0)
Regress the expression of three genes against the medium factor using formula:
>>> kfit = AOV.from_formula('AACS + ACSL6 + ACSS1 ~ C(Medium, OneHot)', data=data) >>> print(kfit.data.exog_names) Index(['Medium[0H]', 'Medium[24H]', 'Medium[48HDIFF]', 'Medium[48HREV]'], dtype='object') >>> print(kfit.data.endog_names) Index(['AACS', 'ACSL6', 'ACSS1'], dtype='object')
- compute_diagnostics(n_trunc=100, n_anchors=None)[source]#
Calculates diagnostics associated with the model and saves them in the diagnostics attribute. The latter is a dictionary containing the following quantities: - Embeddings: projections of the embeddings on the first eigenfunctions of the residual covariance operator. - Predictions: projections of the predictions of the embeddings on the first eigenfunctions of the residual covariance operator. - Residuals: projections of the residuals on the first eigenfunctions of the residual covariance operator.
- Parameters:
- n_truncint, optional
Maximal truncation for projections calculation, the default is 100.
- n_anchorsint, optional
Number of anchors used in the Nystrom method. If None, the value is set at n_trunc.
- plot_diagnostics(trunc=100, diagnostic='residuals', n_trunc=100, n_anchors=None, colormap='viridis', alpha=0.75, legend_fontsize=12, font_family='serif', figsize=None)[source]#
Plots diagnostics associated with the model. The graph contains sublots for each factor, with the projections of either the residuals (diagnostic=’residuals’) or the embeddings (diagnostic=’embeddings’) on the first eigenfunctions of the residual covariance operator, plotted against the projections of the predictions on these eigenfunctions.
- Parameters:
- truncint, optional
Truncation to plot, by default equal to the maximal truncation.
- diagnosticstr, optional
The type of diagnostic to plot against the predictions. The default is ‘residuals’, the alternative is ‘embeddings’.
- n_truncint, optional
Maximal truncation for projections calculation, the default is 100.
- n_anchorsint, optional
Number of anchors used in the Nystrom method. If None, the value is set at n_trunc.
- colormapstr, optional
The name of a matplotlib colormap to be used for different factor levels. The default is ‘viridis’.
- alphafloat, optional
The alpha blending value, between 0 (transparent) and 1 (opaque). The default is 0.5.
- legend_fontsizeint, optional
Legend font size. The default is 15.
- font_familystr, optional
Legend and labels’ font family name accepted by matplotlib (e.g., ‘serif’, ‘sans-serif’, ‘monospace’, ‘fantasy’ or ‘cursive’), the default is ‘serif’.
- figsizetuple, optional
The size of the figure. If not specified, is set to (8 * nb_factors, 6).
- Returns:
- figmatplotlib.figure.Figure
A Figure object of the plot.
- axsnumpy.ndarray of matplotlib.axes._axes.Axes
An Axes object of the plot.
- set_hypotheses(hypotheses=None, by_level=False, test_intercept=False, true_proportions=False)[source]#
Set hypotheses to be tested.
- Parameters:
- hypothesesstr or None or list[tuple]
Hypotheses to be tested. - if str: either ‘pairwise’ (default for OneHot) or ‘one-vs-all’. Recommended options in combination with OneHot encoding. If ‘pairwise’, the levels of a factor are compared one to another in the pairwise way. If ‘one-vs-all’, each level is compared to the factor mean. - if None: produces an identity contrast matrix for each factor. Intended for other coding schemes (e.g. Treatment, Sum, etc). - if list[tuple]: custom hypothesis option. Each element of the list should be a tuple of size 2: (name, contrast_L), where name is a string and contrast_L is a contrast matrix in the form of a torch.tensor (dtype=torch.float64).
- by_levelbool, optional
If False (default), computes the global test. If True, computes the test by level or by a pair of levels.
- test_interceptbool, optional
If True, adds a test for the intercept, which is set as the grand mean (or the actual mean if true_proportions=True) of all the level effects. The default is False. It is unnecessary to add an intercept test manually eith this option if an intercept is present in the design matrix.
- true_proportionsbool, optional
Relevant for the calculation of the factor mean, i.e. if hypotheses=’one-vs-all’ or test_intercept=True. If False (default), the factor mean is the grand mean (mean of means). If True, the true level proportions are taken into account, so the factor mean is the actual global mean of the factor.
- Returns:
- hypotheseslist[tuple]
List of hypotheses to be tested. Each element of the list is a tuple of size 2: (name, contrast_L), where name is a string and contrast_L is a contrast matrix in the form of a torch.tensor.
- project_on_discriminant(K, K_T, D, center=True)[source]#
Computes of embeddings of each observation on the discriminant axes associated with the test. The latter are correspond to the eigenfunctions of the test statistic operator.
- Parameters:
- Ktorch.tensor
The gram matrix.
- K_Ttorch.tensor
A transformed gram matrix, obtained with _compute_K_T.
- Dtorch.tensor
A quantity containing the information on the contrast matrix, obtained with _compute_D.
- centerbool, optional
If True (default), the projections are centered with respect to the factor mean.
- Returns:
- projpandas.DataFrame
Contains projection values, with rows corresponding to observations, and columns to truncations of the residual covariance operator.
- compute_cook_distances(L, n_trunc=100, normalize=True)[source]#
Computes influence measures in the form of Cook’s distances for each observation as well as the corresponding p-values. The p-values are obtained based on the Beta distribution associated with normalized Cook’s distances.
- Parameters:
- Ltorch.tensor
Contrast matrix, defining a statistical test of interest.
- n_truncint, optional
Maximal truncation of the residual covariance operator, the default is 100.
- normalizebool, optional
If True (default), the distances are normalized with respect to the design and contrasts.
- Returns:
- cook_distancespandas.DataFrame
Contains influence measure values, with rows corresponding to observations, and columns to truncations of the residual covariance operator.
- cook_pvaluespandas.DataFrame
Contains p-values associated with the influence measure, with rows corresponding to observations, and columns to truncations of the residual covariance operator.
- correct_pvalues(pvalues, correction='bonferroni', by_level=False)[source]#
Corrects the p-values according to the chosen correction strategy.
- Parameters:
- pvaluespandas.DataFrame
Data frame with p-values to correct. Columns indicate hypotheses that are tested and lines indicate truncation levels.
- correctionstr, optional
Relevent for multiple test comparisons, in particular when by_level=True. If ‘bonferroni’ (default), permorms the Bonferroni correction of the p-values. If ‘BH’, perfoms the Benjamini–Hochberg correction.
- by_levelby_levelbool, optional
If False (default), computes the global test. If True, computes the test by level or by a pair of levels.
- test(hypotheses=None, hypotheses_subset=None, by_level=False, n_trunc=100, correction=None, test_intercept=False, true_proportions=False, center_projections=True, verbose=1, n_anchors=None, f_norm=True, skip_projections_and_cook=False, norm_cook=True)[source]#
Performs kernel hypothesis tests for the given model. Simultaneously calculates projections on the associated discriminant axes as well as influences of observations with respect to the test with their p-values.
- Parameters:
- hypothesesstr or None or list[tuple]
Hypotheses to be tested. - if str: either ‘pairwise’ (default for OneHot) or ‘one-vs-all’. Recommended options in combination with OneHot encoding. - if None: produces an identity contrast matrix for each factor. Intended for other coding schemes (e.g. Treatment, Sum, etc). - if list[tuple]: custom hypothesis option. Each element of the list should be a tuple of size 2: (name, contrast_L), where name is a string and contrast_L is a contrast matrix in the form of a torch.tensor (dtype=torch.float64).
- hypotheses_subsetlist of strings
Names of tests to perform, subset of all the tests in the hypotheses variable (particularly useful with the by_level testing option, when the total number of hypotheses is high and the interest lies in the subset). The default is None, i.e. all hypotheses are tested.
- by_levelbool, optional
If False (default), computes the global test. If True, computes the test by level or by a pair of levels.
- n_truncint, optional
Maximal truncation for statistics calculation, the default is 100.
- correctionstr or None, optional
Relevent for multiple test comparisons, in particular when by_level=True. If ‘bonferroni’, permorms the Bonferroni correction of the p-values. If ‘BH’, perfoms the Benjamini–Hochberg correction. If None (default), the p-values remain uncorrected.
- test_interceptbool, optional
If True, adds a test for the intercept, which is set as the grand mean (or the actual mean if true_proportions=True) of all the level effects. The default is False. It is unnecessary to add an intercept test manually eith this option if an intercept is present in the design matrix.
- true_proportionsbool, optional
Relevant for the calculation of the factor mean, i.e. if hypotheses=’one-vs-all’ or test_intercept=True. If False (default), the factor mean is the grand mean (mean of means). If True, the true level proportions are taken into account, so the factor mean is the actual global mean of the factor.
- center_projectionsbool, optional
If True (default), the projections are centered with respect to the factor mean.
- n_anchorsint, optional
Number of anchors used in the Nystrom method. If None, the value is set at n_trunc.
- verboseint, optional
The higher the verbosity, the more messages keeping track of computations. The default is 0. - < 1: no messages, - 1: progress bar with computation time, - 2: print tested hypothesis’ name, warnings are printed once, - 3: warnings are printed every time they appear.
- f_normbool, optional
If True (default), the test statistic is normalized and asymptotically follows an f-distribution. Otherwise, the original chi-2 version is returned.
- skip_projections_and_cookbool, optional
If False (default), projections on the discriminant axes as well as Cook’s distances will be computed. Set to True to avoid computing them if they are not needed, in order to reduce computation time.
- norm_cookbool, optional
If True (default), the Cook’s distances are normalized with respect to the design and contrasts.
- Returns:
- KernelAOVResults object
See the documentation for KernelAOVResults.
Examples
Importing data and create an intsance of AOV:
>>> import pandas as pd >>> from kaov import AOV >>> url = "https://raw.githubusercontent.com/LMJL-Alea/kAOV/refs/heads/main/Data/reversion_kAOV.csv" >>> data = pd.read_csv(url, index_col=0) >>> kfit = AOV.from_formula('AACS + ACSL6 + ACSS1 ~ C(Medium, OneHot)', data=data)
Test for factor effects:
>>> res = kfit.test() >>> print(res) Kernel Analysis of Variance (trunc. 1): ==================================== ------------------------------------ Factor test | factor stat pval ------------------------------------ | Medium 34.0842 0.0000 ====================================
- class kaov.kaov.KernelAOVResults(hypotheses, stats, projections, cook_distances, cook_pvalues, hypothesis_type, by_level, factor_info)[source]#
Bases:
objectClass implementing Kernel Analysis Of Variance.
- Parameters:
- hypotheseslist
All the consideres hypotheses. Each element of the list represents a hypothesis in the form of a list with two elements: a name and a congtrast matrix associated with the test.
- statsdict
A dictionary with keys corresponding to hypothesis names, and values to instances of pandas.DataFrame with the results of the corresponding tests. Each data frame contains two columns, the first containing the truncated kernel Hotelling-Lawley test statistic values, and the second containing the associated p-values, indexed by truncations of the residual covariance operator used in the calculations.
- projectionsdict
A dictionary with keys corresponding to hypothesis names, and values to instances of pandas.DataFrame with the projections obtained with AOV.project_on_discriminant. In each data frame, rows correspond to observations, and columns to truncations of the residual covariance operator.
- cook_distancesdict
A dictionary with keys corresponding to hypothesis names, and values to instances of pandas.DataFrame with the influence measure values obtained with AOV.compute_cook_distances. In each data frame, rows correspond to observations, and columns to truncations of the residual covariance operator.
- cook_pvaluesdict
A dictionary with keys corresponding to hypothesis names, and values to instances of pandas.DataFrame with the p-values associated with the influence measures obtained with AOV.compute_cook_distances. In each data frame, rows correspond to observations, and columns to truncations of the residual covariance operator.
- hypothesis_typestr or None
Types of hypotheses that were tested. If not None, expected options are ‘pairwise’, ‘one-vs-all’ and ‘custom’.
- by_levelbool, optional
False if global (factor-wise) tests were perfromed, True if tests by level or by a pair of levels were perfromed.
- factor_infodict
_factor_info attribute of the AOV class.
- Attributes:
- hypotheseslist
See Parameters.
- statsdict
See Parameters.
- projectionsdict
See Parameters.
- cook_distancesdict
See Parameters.
- cook_pvaluesdict
See Parameters.
- hypothesis_typestr or None
See Parameters.
- by_levelbool
See Parameters.
- summary(trunc, factor=None)[source]#
Creates a pandas.DataFrame or a distionary with a summary of the test for a given truncation and factor (in the by-level case).
- Parameters:
- truncint
Truncation for which to return the test results.
- factorstr or None
None by default, in which case the results of tests for each factor are returned. If the factor is specified, returns the results of tests on comparisons related with the chosen factor.
- Returns:
- sum_dfpandas.DataFrame or pandas.Series or dict
A data frame with the summary of test results. If not by_level and factor is specified, returns a row of this data frame corresponding to the chosen factor. If by_level and factor is not specified, returns a dictionary with data frames for each factor.
- plot_density(comp=1, tests=None, colormap='viridis', alpha=0.5, legend_fontsize=12, font_family='serif', figsize=None)[source]#
Plots kernel-densities of projections of the embeddings on the chosen discriminant axis, associated with the tests underlying the KernelAOVResults object. Produces separate subplots for each test.
- Parameters:
- compint, optional
Component to plot, i.e. the embeddings are projected on the comp-th eigenfunction.
- testslist of strings or None
List containing names of tests to plot, out of all the tests in the KernelAOVResults object (particularly useful with the by_level testing option). The default is None, i.e. all tests are plotted.
- colormapstr, optional
The name of a matplotlib colormap to be used for different factor levels. The default is ‘viridis’.
- alphafloat, optional
The alpha blending value, between 0 (transparent) and 1 (opaque). The default is 0.5.
- legend_fontsizeint, optional
Legend font size. The default is 15.
- font_familystr, optional
Legend and labels’ font family name accepted by matplotlib (e.g., ‘serif’, ‘sans-serif’, ‘monospace’, ‘fantasy’ or ‘cursive’), the default is ‘serif’.
- figsizetuple, optional
The size of the figure. If not specified, is set to (8 * nb_factors, 6).
- Returns:
- figmatplotlib.figure.Figure
A Figure object of the plot.
- axsnumpy.ndarray of matplotlib.axes._axes.Axes
An Axes object of the plot.
- plot_mean_embedding_projections(comp1=1, comp2=2, tests=None, figsize=None, ylim=None, xlim=None, alpha=1, s=50, marker='o', colors=None, colormap='viridis', font_family='serif', legend=True, legend_fontsize=15, figtitle=None)[source]#
Plots projections of mean embeddings of relevat groups on the chosen discriminant axis, associated with the tests underlying the KernelAOVResults object. Produces separate subplots for each test.
- Parameters:
- comp1int, optional
Component x of the plot, i.e. the mean embeddings are projected on the comp1-th eigenfunction.
- comp2int, optional
Component y of the plot, i.e. the mean embeddings are projected on the comp2-th eigenfunction.
- testslist of strings or None
List containing names of tests to plot, out of all the tests in the KernelAOVResults object (particularly useful with the by_level testing option). The default is None, i.e. all tests are plotted.
- figsizetuple, optional
The size of the figure. If not specified, is set to (8 * nb_factors, 6).
- alphafloat, optional
The alpha blending value, between 0 (transparent) and 1 (opaque). The default is 1.
- sint or dict
Marker sizes of the mean embedding projections. The default is 50. If int, same sizes for all mean embeddings. To pass different values for different tests and levels, pass a dictionary with keys corresponding to test names, and values that are dictionaries with keys corresponding to the levels of the test and values that are marker size integers.
- markerstr or dict
Marker styles of the mean embedding projections. The default is ‘o’. If string, same marker for all mean embeddings. To pass different values for different tests and levels, pass a dictionary with keys corresponding to test names, and values that are dictionaries with keys corresponding to the levels of the test and values that are marker style strings.
- colorsNone or dict
Colors of the mean embedding projections. If None (default), colors are chosen from the specified colormap. To customize, pass a dictionary with keys corresponding to test names, and values that are dictionaries with keys corresponding to the levels of the test and values that are colors.
- colormapstr, optional
The name of a matplotlib colormap to be used for different factor levels (if colors are not provided). The default is ‘viridis’.
- font_familystr, optional
Legend and labels’ font family name accepted by matplotlib (e.g., ‘serif’, ‘sans-serif’, ‘monospace’, ‘fantasy’ or ‘cursive’), the default is ‘serif’.
- legendbool, optional
If True (default), legend is plotted automatically.
- legend_fontsizeint, optional
Legend font size. The default is 15.
- figtitlestr, optional
The title of the figure.
- Returns:
- figmatplotlib.figure.Figure
A Figure object of the plot.
- axsnumpy.ndarray of matplotlib.axes._axes.Axes
An Axes object of the plot.
- plot_influence(trunc=1, comp=1, tests=None, marker='o', colors=None, colormap='viridis', font_family='serif', legend=True, alpha=0.5, legend_fontsize=12, figsize=None)[source]#
Plots influences (Cook’s distances) of the embeddings, associated with the tests underlying the KernelAOVResults object, against their projections on the chosen discriminant axis. Produces separate subplots for each test.
- Parameters:
- truncint, optional
Truncation of the resdual covariance operator used for the Cook’s distance calculation.
- compint, optional
Component of the projections, i.e. the embeddings are projected on the comp-th eigenfunction.
- testslist of strings
List containing a list of tests to plot, out of all the tests in the KernelAOVResults object (particularly useful with the by_level testing option).
- markerstr or dict
Marker styles of the embedding projections. The default is ‘o’. If string, same marker for all mean embeddings. To pass different values for different tests and levels, pass a dictionary with keys corresponding to test names, and values that are dictionaries with keys corresponding to the levels of the test and values that are marker style strings.
- colorsNone or dict
Colors of the mean embedding projections. If None (default), colors are chosen from the specified colormap. To customize, pass a dictionary with keys corresponding to test names, and values that are dictionaries with keys corresponding to the levels of the test and values that are colors.
- colormapstr, optional
The name of a matplotlib colormap to be used for different factor levels (if colors are not provided). The default is ‘viridis’.
- alphafloat, optional
The alpha blending value, between 0 (transparent) and 1 (opaque). The default is 0.5.
- legend_fontsizeint, optional
Legend font size. The default is 15.
- font_familystr, optional
Legend and labels’ font family name accepted by matplotlib (e.g., ‘serif’, ‘sans-serif’, ‘monospace’, ‘fantasy’ or ‘cursive’), the default is ‘serif’.
- figsizetuple, optional
The size of the figure. If not specified, is set to (8 * nb_factors, 6).
- Returns:
- figmatplotlib.figure.Figure
A Figure object of the plot.
- axsnumpy.ndarray of matplotlib.axes._axes.Axes
An Axes object of the plot.
- get_projections(n_comp=None, factor=None, hypothesis=None)[source]#
Creates a pandas.DataFrame with the discriminant axis projections associated with the test for a given number of components, factor (for all tests of type ‘pairwise’ and ‘one-vs-all’) and hypothesis (in the by-level cases and for user-specified tests).
- Parameters:
- n_compint or None
Number of discriminant axis components for which to return the results. If None (default), projections on all available components are returned.
- factorstr or None
None by default (acceptable for user-specified tests) only. A factor has to be specified in all other cases, then returns the results of tests on comparisons related with the chosen factor. Factor names are keys of the _factor_info attribute of the AOV class.
- hypothesisstr or None
None by default, which is acceptable if the test is global or of type ‘one-vs-all’. In the by-level pairwise or custom test cases a hypothesis has to be specified. The list of possible hypothesis names is accessible through the hypotheses attribute of the KernelAOVResults class (first element of each tuple).
- Returns:
- sum_dfpandas.DataFrame
A data frame with the summary of projections.
- get_cook(n_trunc=None, factor=None, hypothesis=None)[source]#
Creates a pandas.DataFrame with Cook’s distances and their p-values associated with the test for given truncations, factor (for all tests of type ‘pairwise’ and ‘one-vs-all’) and hypothesis (in the by-level cases and for user-specified tests).
- Parameters:
- n_truncint or None
Number of truncations for which to return the results. If None (default), Cook’s distances are returned for all available truncations.
- factorstr or None
None by default (acceptable for user-specified tests) only. A factor has to be specified in all other cases, then returns the results of tests on comparisons related with the chosen factor. Factor names are keys of the _factor_info attribute of the AOV class.
- hypothesisstr or None
None by default, which is acceptable if the test is global or of type ‘one-vs-all’. In the by-level pairwise or custom test cases a hypothesis has to be specified. The list of possible hypothesis names is accessible through the hypotheses attribute of the KernelAOVResults class (first element of each tuple).
- Returns:
- sum_dfpandas.DataFrame
A data frame with the summary of projections.