Skip to content

Built-in Plots

Performance Plots

ApproximationRatioPlot

Bases: MetricBarPlot

Bar chart of the approximation ratio per algorithm.

The approximation ratio compares an algorithm's objective value against the known optimum, so 100% - marked by the reference line - means the optimum was reached. One bar per algorithm, aggregated over every model - the mean by default, see :attr:aggregation - with the spread across those models as the error bar.

Requires the ApproximationRatio metric.

Every display option is inherited and can be set when the plot is constructed, e.g. ApproximationRatioPlot(annotate=False, file_formats=("pgf", "png")). BarPlot documents the colours, error bars, value annotations and grouping by a feature; SeabornPlot the figure size and the output formats.

Attributes:

  • figure_filename (str) –

    Stem of the written figure files, by default "approximation_ratio".

Examples:

>>> bench.add_metric(name="approx_ratio", metric=ApproximationRatio())
>>> bench.add_plot(name="avg_approx", plot=ApproximationRatioPlot())
See Also

ApproximationRatioVsVarNumberPlot : The same ratio against model size.

theme: Theme | None = Theme() class-attribute instance-attribute

The seaborn theme the figure is drawn under, and the gridlines behind the marks.

A figure is read against its axis, so it is drawn with the lines that make that possible unless asked otherwise. None takes the theme and the grid away.

missing: Missing = Missing() class-attribute instance-attribute

What becomes of the values the plot cannot draw, and how it says they were there.

option_bundles: dict[str, type[OptionBundle]] = {'style': PlotStyle} class-attribute

Constructor arguments that configure several fields at once, and the bundle each takes.

A style is spread over the bundle fields rather than stored, so a benchmark can hand the same look to every plot while each keeps what it says itself. Applied least specific first: the shared style, then a bundle passed to the plot, then a flat option.

x: Dimension = AlgorithmDimension() class-attribute instance-attribute

What the bars are: one per value of this dimension, and its title on the axis.

aggregation: Aggregation = Aggregation.MEAN class-attribute instance-attribute

Aggregation applied to the values of an x category, by default their mean.

errorbars: ErrorBars | None = ErrorBars() class-attribute instance-attribute

The error bars drawn on top of the bars: what they show, their colour and caps.

None draws none, the same as ErrorBars(spec=None).

annotation: Annotation | None = None class-attribute instance-attribute

The values written above the bars - how they are formatted and how large.

None, the default, writes none: a bar chart is read off its axis, and a number above every bar is worth its clutter only when the exact value is the point. Pass an Annotation to turn them on, empty for the defaults.

grouping: Dimension | None = None class-attribute instance-attribute

What splits each bar into a group of bars.

One of the groupers - ModelDimension, AlgorithmDimension, FeatureDimension, ParameterDimension - or None, which leaves the bars ungrouped.

metric_cls: MetricClass property

The metric this plot reads, taken from its @plot(...) declaration.

Returns:

  • MetricClass –

    The first declared metric class.

Raises:

  • PlotMetricUndeclaredError –

    If the plot declares no metric, so there is nothing to read.

run(benchmark_results: BenchmarkResultContainer, save_dir: str | None = None) -> None

Generate plot output from benchmark results.

Parameters:

  • benchmark_results (BenchmarkResultContainer) –

    Aggregated benchmark data consumed by the plot implementation.

  • save_dir (str | None, default: None ) –

    Directory to save the figure into, by default None.

draw_into(axes: Axes, benchmark_results: BenchmarkResultContainer) -> None

Run this plot onto an existing axes rather than into a figure of its own.

Everything the plot draws goes through the pyplot state, so making that axes the current one is enough to redirect it. The figure, the files, and the window stay the caller's business - which is what lets several plots share one figure, e.g. the summary grid.

Parameters:

  • axes (Axes) –

    The axes to draw on.

  • benchmark_results (BenchmarkResultContainer) –

    Aggregated benchmark data handed to :meth:run.

resolve_missing(df: pd.DataFrame, column: str, *, by: str | None = None, within: str | None = None) -> tuple[pd.DataFrame, dict[tuple[str, str], int]]

Return the drawable rows and how many of them were not, per category.

A metric with nothing to report says so with a None or an infinity - a time to solution of a run that never reached the optimum is the usual one, since the expected time to something that did not happen is unbounded. Neither is a height a bar can have, and leaving them in poisons the aggregate: one infinity turns the mean of an algorithm into an infinity, and a missing value silently shortens it.

What happens to them is :attr:missing, and by default it is nothing: the plot raises rather than quietly showing a mean over fewer models than it claims. Asked to carry on, it either leaves them out or fills them from the values that could be drawn - Missing(policy="max") puts them just past the tallest bar, which is where "worse than everything here" belongs. Either way they are counted and a warning is logged: a bar resting on half its models is a different statement from one resting on all of them, and that is not visible in the bar itself.

Parameters:

  • df (DataFrame) –

    The plotting data.

  • column (str) –

    Column holding the plotted value.

  • by (str | None, default: None ) –

    Column whose categories the missing values are counted per, e.g. the x-axis of a bar plot. Without one they are counted under "".

  • within (str | None, default: None ) –

    Column that splits those categories further, e.g. the grouping of a bar plot. Counting per group is what lets a figure mark the one bar of a group that lost values rather than the whole category it sits in.

Returns:

  • tuple[DataFrame, dict[tuple[str, str], int]] –

    The rows to draw, and the number of missing values per category and group - the group is "" where there is none. The mapping is empty when nothing was missing.

Raises:

  • PlotMissingValuesError –

    If values are missing and the policy is "raise".

place_legend(axes: Axes, handles: list[Any] | None = None, labels: list[str] | None = None) -> None

Put the legend beside the axes, whatever drew it.

Outside the axes for a figure of its own: a legend inside sits on top of the data, and which corner is free depends on the run rather than on the plot - the figure would move its own key around as the numbers change. Beside it, the key is always in the same place and covers nothing.

A panel of someone else's figure is the exception. The room beside it belongs to the panel next to it, so a key anchored there is drawn over a neighbour rather than over the data; inside the panel it stays within the space the plot was given.

Parameters:

  • axes (Axes) –

    The axes the plot was drawn on.

  • handles (list[Any] | None, default: None ) –

    Legend handles, by default the ones already on the axes.

  • labels (list[str] | None, default: None ) –

    Their labels, by default the ones already on the axes.

note_missing(handles: list[Any], labels: list[str], missing: dict[tuple[str, str], int]) -> None

Add the legend entry that says how many values the figure could not draw.

What a plot can say beyond that depends on what it draws. A bar has a slot of its own to put a cross under, so BarPlot marks the categories themselves; a point in a cloud or a step of a sweep has no slot, and the count in the key is the whole statement there - enough that a filled value is not read as a measured one.

Parameters:

  • handles (list[Any]) –

    Legend handles, extended in place.

  • labels (list[str]) –

    Their labels, extended in place alongside handles.

  • missing (dict[tuple[str, str], int]) –

    Number of missing values per category and group, as counted by :meth:resolve_missing.

apply_theme() -> None

Install the seaborn theme this plot is drawn under, unless it has none.

The theme is matplotlib's global state rather than a property of one figure, so it is installed before the figure is built and left in place afterwards: a benchmark themes its plots by handing every one of them the same Theme, not by each plot putting the previous look back.

apply_grid(axes: Axes) -> None

Draw the gridlines the theme asks for, behind everything else on axes.

Parameters:

  • axes (Axes) –

    The axes the plot was drawn on.

setup_figure() -> None

Create a matplotlib figure, unless the plot is drawing into a shared axes.

save_figure(save_dir: str) -> list[Path]

Write the current figure to save_dir once per configured file format.

Parameters:

  • save_dir (str) –

    Directory to save the figure into. Created if it does not exist.

Returns:

  • list[Path] –

    Paths that were written successfully.

finalize_plot(xlabel: str, ylabel: str, title: str, ylim: tuple[float, float] | None = None, x_rotation: int = 45, save_dir: str | None = None) -> None

Apply common axis labels, title, limits, and display behavior.

Parameters:

  • xlabel (str) –

    Label for the x-axis.

  • ylabel (str) –

    Label for the y-axis.

  • title (str) –

    Plot title.

  • ylim (tuple[float, float] | None, default: None ) –

    Lower and upper y-axis limits, by default None.

  • x_rotation (int, default: 45 ) –

    Rotation angle for x-axis tick labels, by default 45.

  • save_dir (str | None, default: None ) –

    Directory to save the figure into, by default None.

apply_grouping(benchmark_results: BenchmarkResultContainer, rows: list[dict[str, Any]]) -> dict[str, Any]

Split rows into groups along :attr:grouping.

What that means is the grouper's business - a column of the plotted data, a value looked up per model, or a setting the algorithms were configured with - and so is deciding that it does not apply, in which case the bars stay ungrouped.

Parameters:

  • benchmark_results (BenchmarkResultContainer) –

    Benchmark data the feature results and algorithm configurations are read from.

  • rows (list[dict[str, Any]]) –

    Row-oriented plot data, annotated - and, where a grouping applies to only part of the data, reduced - in place.

Returns:

  • dict[str, Any] –

    Keyword arguments to forward to :meth:create. Empty when no grouping applies, so call sites can splat it unconditionally.

draw(*, benchmark_results: BenchmarkResultContainer, rows: list[dict[str, Any]], save_dir: str | None = None, **overrides: Any) -> None

Group rows and draw them with the display configuration of this plot.

This is what turns the declared fields - :attr:x, :attr:title, :attr:hline and the rest - into a :meth:create call, so a subclass only has to say which rows it plots. Doing it in one place is also what keeps :attr:group_by working for every bar plot rather than for those that remember to apply it.

Parameters:

  • benchmark_results (BenchmarkResultContainer) –

    Benchmark data, used to look up the groups of a feature :attr:group_by.

  • rows (list[dict[str, Any]]) –

    Row-oriented plot data.

  • save_dir (str | None, default: None ) –

    Directory to save the figure into, by default None.

  • **overrides (Any, default: {} ) –

    Keyword arguments forwarded to :meth:create, overriding the fields.

transform_rows(rows: list[dict[str, Any]], x: str | None, group: str | None) -> list[dict[str, Any]]

Return the rows to plot, by default the rows as they are.

A subclass that has to reduce its rows before they are drawn - pooling counts into a single ratio, say - overrides this rather than :meth:run, so it keeps the shared grouping and display handling. It is told what the bars and the groups turned out to be, since that is what a row has to keep to stay one of them.

Parameters:

  • rows (list[dict[str, Any]]) –

    Row-oriented plot data, already annotated with the dimensions' columns.

  • x (str | None) –

    Column the bars are drawn per, or None when the plot has no rows.

  • group (str | None) –

    Column the bars are split by, or None when they are ungrouped.

Returns:

create(*, rows: list[dict[str, Any]], xlabel: str, ylabel: str, title: str, x: str = 'x', y: str = 'y', aggregation: Aggregation = Aggregation.MEAN, errorbar: ErrorBar | str = AUTO_ERRORBAR, hue: str | None = None, hline: float | None = None, hline_label: str | None = None, hcolor: str = REFERENCE_LINE_COLOUR, baseline: float | None = None, ylim: tuple[float, float] | None = None, legend: bool = False, save_dir: str | None = None, **kwargs: Any) -> None

Create a bar plot from row-oriented data.

Parameters:

  • rows (dict[str, Any]) –

    Row-oriented mapping used to construct the plotting DataFrame.

  • xlabel (str) –

    Label for the x-axis.

  • ylabel (str) –

    Label for the y-axis.

  • title (str) –

    Plot title.

  • x (str, default: 'x' ) –

    Column name mapped to the x-axis, by default "x".

  • y (str, default: 'y' ) –

    Column name mapped to the y-axis, by default "y".

  • aggregation (Aggregation, default: MEAN ) –

    Aggregation strategy applied by seaborn, by default Aggregation.MEAN.

  • errorbar (ErrorBar | str, default: AUTO_ERRORBAR ) –

    Seaborn error bar specification ("sd", ("ci", 95), None to disable). By default "auto", which takes the error bar from aggregation: the spread of the samples for means, none for min/max.

  • hue (str | None, default: None ) –

    Optional grouping column for grouped bars, by default None.

  • hline (float | None, default: None ) –

    Optional horizontal reference line value, by default None.

  • hline_label (str | None, default: None ) –

    Legend label for the horizontal reference line, by default None.

  • hcolor (str, default: REFERENCE_LINE_COLOUR ) –

    Colour of the horizontal reference line, by default black.

  • baseline (float | None, default: None ) –

    Height of a solid black baseline marking where the bars start, by default None. Unlike hline it carries no label and stays out of the legend - it says where zero is, it does not name a target.

  • ylim (tuple[float, float] | None, default: None ) –

    Lower and upper y-axis limits, by default None.

  • legend (bool, default: False ) –

    Whether seaborn should create a legend for hue groups, by default False.

  • save_dir (str | None, default: None ) –

    Directory to save the figure into, by default None.

  • **kwargs (Any, default: {} ) –

    Additional keyword arguments forwarded to :func:seaborn.barplot. They override the defaults computed here, so anything seaborn understands (palette, saturation, capsize, err_kws, ...) can be tuned from the call site.

annotation_text(value: float) -> str

Return the text written above a bar of value.

Applies :attr:annotate_format, unless :attr:annotate_max_decimals allows the value to be written as a plain decimal instead of in scientific notation. A value read off a percent axis is written as a percentage, so the annotation says the same thing as the axis it stands on - unless a format was asked for, which wins.

Parameters:

  • value (float) –

    The aggregated value of one bar.

Returns:

  • str –

    The annotation, e.g. "0.000057" rather than "5.67e-05".

value(metric_result: MetricResult) -> float

Return the number a single metric result contributes.

Parameters:

  • metric_result (MetricResult) –

    One result of :attr:metric_cls.

Returns:

  • float –

    The value plotted for this result, by default the attribute :attr:y names.

rows(benchmark_results: BenchmarkResultContainer) -> list[dict[str, Any]]

Return one row per model and algorithm.

Parameters:

Returns:

  • list[dict[str, Any]] –

    Rows carrying the algorithm, the model, and the plotted value under :attr:y.

BestSolutionFoundRatioPlot

Bases: MetricBarPlot

Bar chart of the best-solution-found ratio per algorithm.

The share of samples that reached the best solution, averaged over every model, with the spread across those models as the error bar. Higher is better.

Requires the BestSolutionFoundRatio metric.

Every display option is inherited and can be set when the plot is constructed, e.g. BestSolutionFoundRatioPlot(annotate=False, file_formats=("pgf", "png")). BarPlot documents the colours, error bars, value annotations and grouping by a feature; SeabornPlot the figure size and the output formats.

Attributes:

  • figure_filename (str) –

    Stem of the written figure files, by default "best_solution_found_ratio".

Examples:

>>> bench.add_metric(name="bsf_ratio", metric=BestSolutionFoundRatio())
>>> bench.add_plot(name="avg_bsfr", plot=BestSolutionFoundRatioPlot())

baseline: float | None = 0.0 class-attribute instance-attribute

Zero line the bars stand on.

A wide error bar can reach below zero, which leaves the bars floating in the axes without it - and where zero is, is what tells a ratio of 1.2 from one of 0.2.

theme: Theme | None = Theme() class-attribute instance-attribute

The seaborn theme the figure is drawn under, and the gridlines behind the marks.

A figure is read against its axis, so it is drawn with the lines that make that possible unless asked otherwise. None takes the theme and the grid away.

missing: Missing = Missing() class-attribute instance-attribute

What becomes of the values the plot cannot draw, and how it says they were there.

option_bundles: dict[str, type[OptionBundle]] = {'style': PlotStyle} class-attribute

Constructor arguments that configure several fields at once, and the bundle each takes.

A style is spread over the bundle fields rather than stored, so a benchmark can hand the same look to every plot while each keeps what it says itself. Applied least specific first: the shared style, then a bundle passed to the plot, then a flat option.

x: Dimension = AlgorithmDimension() class-attribute instance-attribute

What the bars are: one per value of this dimension, and its title on the axis.

aggregation: Aggregation = Aggregation.MEAN class-attribute instance-attribute

Aggregation applied to the values of an x category, by default their mean.

errorbars: ErrorBars | None = ErrorBars() class-attribute instance-attribute

The error bars drawn on top of the bars: what they show, their colour and caps.

None draws none, the same as ErrorBars(spec=None).

annotation: Annotation | None = None class-attribute instance-attribute

The values written above the bars - how they are formatted and how large.

None, the default, writes none: a bar chart is read off its axis, and a number above every bar is worth its clutter only when the exact value is the point. Pass an Annotation to turn them on, empty for the defaults.

grouping: Dimension | None = None class-attribute instance-attribute

What splits each bar into a group of bars.

One of the groupers - ModelDimension, AlgorithmDimension, FeatureDimension, ParameterDimension - or None, which leaves the bars ungrouped.

metric_cls: MetricClass property

The metric this plot reads, taken from its @plot(...) declaration.

Returns:

  • MetricClass –

    The first declared metric class.

Raises:

  • PlotMetricUndeclaredError –

    If the plot declares no metric, so there is nothing to read.

run(benchmark_results: BenchmarkResultContainer, save_dir: str | None = None) -> None

Generate plot output from benchmark results.

Parameters:

  • benchmark_results (BenchmarkResultContainer) –

    Aggregated benchmark data consumed by the plot implementation.

  • save_dir (str | None, default: None ) –

    Directory to save the figure into, by default None.

draw_into(axes: Axes, benchmark_results: BenchmarkResultContainer) -> None

Run this plot onto an existing axes rather than into a figure of its own.

Everything the plot draws goes through the pyplot state, so making that axes the current one is enough to redirect it. The figure, the files, and the window stay the caller's business - which is what lets several plots share one figure, e.g. the summary grid.

Parameters:

  • axes (Axes) –

    The axes to draw on.

  • benchmark_results (BenchmarkResultContainer) –

    Aggregated benchmark data handed to :meth:run.

resolve_missing(df: pd.DataFrame, column: str, *, by: str | None = None, within: str | None = None) -> tuple[pd.DataFrame, dict[tuple[str, str], int]]

Return the drawable rows and how many of them were not, per category.

A metric with nothing to report says so with a None or an infinity - a time to solution of a run that never reached the optimum is the usual one, since the expected time to something that did not happen is unbounded. Neither is a height a bar can have, and leaving them in poisons the aggregate: one infinity turns the mean of an algorithm into an infinity, and a missing value silently shortens it.

What happens to them is :attr:missing, and by default it is nothing: the plot raises rather than quietly showing a mean over fewer models than it claims. Asked to carry on, it either leaves them out or fills them from the values that could be drawn - Missing(policy="max") puts them just past the tallest bar, which is where "worse than everything here" belongs. Either way they are counted and a warning is logged: a bar resting on half its models is a different statement from one resting on all of them, and that is not visible in the bar itself.

Parameters:

  • df (DataFrame) –

    The plotting data.

  • column (str) –

    Column holding the plotted value.

  • by (str | None, default: None ) –

    Column whose categories the missing values are counted per, e.g. the x-axis of a bar plot. Without one they are counted under "".

  • within (str | None, default: None ) –

    Column that splits those categories further, e.g. the grouping of a bar plot. Counting per group is what lets a figure mark the one bar of a group that lost values rather than the whole category it sits in.

Returns:

  • tuple[DataFrame, dict[tuple[str, str], int]] –

    The rows to draw, and the number of missing values per category and group - the group is "" where there is none. The mapping is empty when nothing was missing.

Raises:

  • PlotMissingValuesError –

    If values are missing and the policy is "raise".

place_legend(axes: Axes, handles: list[Any] | None = None, labels: list[str] | None = None) -> None

Put the legend beside the axes, whatever drew it.

Outside the axes for a figure of its own: a legend inside sits on top of the data, and which corner is free depends on the run rather than on the plot - the figure would move its own key around as the numbers change. Beside it, the key is always in the same place and covers nothing.

A panel of someone else's figure is the exception. The room beside it belongs to the panel next to it, so a key anchored there is drawn over a neighbour rather than over the data; inside the panel it stays within the space the plot was given.

Parameters:

  • axes (Axes) –

    The axes the plot was drawn on.

  • handles (list[Any] | None, default: None ) –

    Legend handles, by default the ones already on the axes.

  • labels (list[str] | None, default: None ) –

    Their labels, by default the ones already on the axes.

note_missing(handles: list[Any], labels: list[str], missing: dict[tuple[str, str], int]) -> None

Add the legend entry that says how many values the figure could not draw.

What a plot can say beyond that depends on what it draws. A bar has a slot of its own to put a cross under, so BarPlot marks the categories themselves; a point in a cloud or a step of a sweep has no slot, and the count in the key is the whole statement there - enough that a filled value is not read as a measured one.

Parameters:

  • handles (list[Any]) –

    Legend handles, extended in place.

  • labels (list[str]) –

    Their labels, extended in place alongside handles.

  • missing (dict[tuple[str, str], int]) –

    Number of missing values per category and group, as counted by :meth:resolve_missing.

apply_theme() -> None

Install the seaborn theme this plot is drawn under, unless it has none.

The theme is matplotlib's global state rather than a property of one figure, so it is installed before the figure is built and left in place afterwards: a benchmark themes its plots by handing every one of them the same Theme, not by each plot putting the previous look back.

apply_grid(axes: Axes) -> None

Draw the gridlines the theme asks for, behind everything else on axes.

Parameters:

  • axes (Axes) –

    The axes the plot was drawn on.

setup_figure() -> None

Create a matplotlib figure, unless the plot is drawing into a shared axes.

save_figure(save_dir: str) -> list[Path]

Write the current figure to save_dir once per configured file format.

Parameters:

  • save_dir (str) –

    Directory to save the figure into. Created if it does not exist.

Returns:

  • list[Path] –

    Paths that were written successfully.

finalize_plot(xlabel: str, ylabel: str, title: str, ylim: tuple[float, float] | None = None, x_rotation: int = 45, save_dir: str | None = None) -> None

Apply common axis labels, title, limits, and display behavior.

Parameters:

  • xlabel (str) –

    Label for the x-axis.

  • ylabel (str) –

    Label for the y-axis.

  • title (str) –

    Plot title.

  • ylim (tuple[float, float] | None, default: None ) –

    Lower and upper y-axis limits, by default None.

  • x_rotation (int, default: 45 ) –

    Rotation angle for x-axis tick labels, by default 45.

  • save_dir (str | None, default: None ) –

    Directory to save the figure into, by default None.

apply_grouping(benchmark_results: BenchmarkResultContainer, rows: list[dict[str, Any]]) -> dict[str, Any]

Split rows into groups along :attr:grouping.

What that means is the grouper's business - a column of the plotted data, a value looked up per model, or a setting the algorithms were configured with - and so is deciding that it does not apply, in which case the bars stay ungrouped.

Parameters:

  • benchmark_results (BenchmarkResultContainer) –

    Benchmark data the feature results and algorithm configurations are read from.

  • rows (list[dict[str, Any]]) –

    Row-oriented plot data, annotated - and, where a grouping applies to only part of the data, reduced - in place.

Returns:

  • dict[str, Any] –

    Keyword arguments to forward to :meth:create. Empty when no grouping applies, so call sites can splat it unconditionally.

draw(*, benchmark_results: BenchmarkResultContainer, rows: list[dict[str, Any]], save_dir: str | None = None, **overrides: Any) -> None

Group rows and draw them with the display configuration of this plot.

This is what turns the declared fields - :attr:x, :attr:title, :attr:hline and the rest - into a :meth:create call, so a subclass only has to say which rows it plots. Doing it in one place is also what keeps :attr:group_by working for every bar plot rather than for those that remember to apply it.

Parameters:

  • benchmark_results (BenchmarkResultContainer) –

    Benchmark data, used to look up the groups of a feature :attr:group_by.

  • rows (list[dict[str, Any]]) –

    Row-oriented plot data.

  • save_dir (str | None, default: None ) –

    Directory to save the figure into, by default None.

  • **overrides (Any, default: {} ) –

    Keyword arguments forwarded to :meth:create, overriding the fields.

transform_rows(rows: list[dict[str, Any]], x: str | None, group: str | None) -> list[dict[str, Any]]

Return the rows to plot, by default the rows as they are.

A subclass that has to reduce its rows before they are drawn - pooling counts into a single ratio, say - overrides this rather than :meth:run, so it keeps the shared grouping and display handling. It is told what the bars and the groups turned out to be, since that is what a row has to keep to stay one of them.

Parameters:

  • rows (list[dict[str, Any]]) –

    Row-oriented plot data, already annotated with the dimensions' columns.

  • x (str | None) –

    Column the bars are drawn per, or None when the plot has no rows.

  • group (str | None) –

    Column the bars are split by, or None when they are ungrouped.

Returns:

create(*, rows: list[dict[str, Any]], xlabel: str, ylabel: str, title: str, x: str = 'x', y: str = 'y', aggregation: Aggregation = Aggregation.MEAN, errorbar: ErrorBar | str = AUTO_ERRORBAR, hue: str | None = None, hline: float | None = None, hline_label: str | None = None, hcolor: str = REFERENCE_LINE_COLOUR, baseline: float | None = None, ylim: tuple[float, float] | None = None, legend: bool = False, save_dir: str | None = None, **kwargs: Any) -> None

Create a bar plot from row-oriented data.

Parameters:

  • rows (dict[str, Any]) –

    Row-oriented mapping used to construct the plotting DataFrame.

  • xlabel (str) –

    Label for the x-axis.

  • ylabel (str) –

    Label for the y-axis.

  • title (str) –

    Plot title.

  • x (str, default: 'x' ) –

    Column name mapped to the x-axis, by default "x".

  • y (str, default: 'y' ) –

    Column name mapped to the y-axis, by default "y".

  • aggregation (Aggregation, default: MEAN ) –

    Aggregation strategy applied by seaborn, by default Aggregation.MEAN.

  • errorbar (ErrorBar | str, default: AUTO_ERRORBAR ) –

    Seaborn error bar specification ("sd", ("ci", 95), None to disable). By default "auto", which takes the error bar from aggregation: the spread of the samples for means, none for min/max.

  • hue (str | None, default: None ) –

    Optional grouping column for grouped bars, by default None.

  • hline (float | None, default: None ) –

    Optional horizontal reference line value, by default None.

  • hline_label (str | None, default: None ) –

    Legend label for the horizontal reference line, by default None.

  • hcolor (str, default: REFERENCE_LINE_COLOUR ) –

    Colour of the horizontal reference line, by default black.

  • baseline (float | None, default: None ) –

    Height of a solid black baseline marking where the bars start, by default None. Unlike hline it carries no label and stays out of the legend - it says where zero is, it does not name a target.

  • ylim (tuple[float, float] | None, default: None ) –

    Lower and upper y-axis limits, by default None.

  • legend (bool, default: False ) –

    Whether seaborn should create a legend for hue groups, by default False.

  • save_dir (str | None, default: None ) –

    Directory to save the figure into, by default None.

  • **kwargs (Any, default: {} ) –

    Additional keyword arguments forwarded to :func:seaborn.barplot. They override the defaults computed here, so anything seaborn understands (palette, saturation, capsize, err_kws, ...) can be tuned from the call site.

annotation_text(value: float) -> str

Return the text written above a bar of value.

Applies :attr:annotate_format, unless :attr:annotate_max_decimals allows the value to be written as a plain decimal instead of in scientific notation. A value read off a percent axis is written as a percentage, so the annotation says the same thing as the axis it stands on - unless a format was asked for, which wins.

Parameters:

  • value (float) –

    The aggregated value of one bar.

Returns:

  • str –

    The annotation, e.g. "0.000057" rather than "5.67e-05".

value(metric_result: MetricResult) -> float

Return the number a single metric result contributes.

Parameters:

  • metric_result (MetricResult) –

    One result of :attr:metric_cls.

Returns:

  • float –

    The value plotted for this result, by default the attribute :attr:y names.

rows(benchmark_results: BenchmarkResultContainer) -> list[dict[str, Any]]

Return one row per model and algorithm.

Parameters:

Returns:

  • list[dict[str, Any]] –

    Rows carrying the algorithm, the model, and the plotted value under :attr:y.

FeasibilityRatioPlot

Bases: MetricBarPlot

Bar chart of the share of feasible samples per algorithm.

Counts how many of the samples an algorithm returned satisfy all constraints, averaged over every model, with the spread across those models as the error bar. 1.0 means every sample was feasible.

Requires the FeasibilityRatio metric.

Every display option is inherited and can be set when the plot is constructed, e.g. FeasibilityRatioPlot(annotate=False, file_formats=("pgf", "png")). BarPlot documents the colours, error bars, value annotations and grouping by a feature; SeabornPlot the figure size and the output formats.

Attributes:

  • figure_filename (str) –

    Stem of the written figure files, by default "feasibility_ratio".

Examples:

>>> bench.add_metric(name="feasibility", metric=FeasibilityRatio())
>>> bench.add_plot(name="avg_feasibility", plot=FeasibilityRatioPlot())
See Also

FeasibleSolutionFoundPlot : Whether any feasible sample was found, per model.

theme: Theme | None = Theme() class-attribute instance-attribute

The seaborn theme the figure is drawn under, and the gridlines behind the marks.

A figure is read against its axis, so it is drawn with the lines that make that possible unless asked otherwise. None takes the theme and the grid away.

missing: Missing = Missing() class-attribute instance-attribute

What becomes of the values the plot cannot draw, and how it says they were there.

option_bundles: dict[str, type[OptionBundle]] = {'style': PlotStyle} class-attribute

Constructor arguments that configure several fields at once, and the bundle each takes.

A style is spread over the bundle fields rather than stored, so a benchmark can hand the same look to every plot while each keeps what it says itself. Applied least specific first: the shared style, then a bundle passed to the plot, then a flat option.

x: Dimension = AlgorithmDimension() class-attribute instance-attribute

What the bars are: one per value of this dimension, and its title on the axis.

aggregation: Aggregation = Aggregation.MEAN class-attribute instance-attribute

Aggregation applied to the values of an x category, by default their mean.

errorbars: ErrorBars | None = ErrorBars() class-attribute instance-attribute

The error bars drawn on top of the bars: what they show, their colour and caps.

None draws none, the same as ErrorBars(spec=None).

annotation: Annotation | None = None class-attribute instance-attribute

The values written above the bars - how they are formatted and how large.

None, the default, writes none: a bar chart is read off its axis, and a number above every bar is worth its clutter only when the exact value is the point. Pass an Annotation to turn them on, empty for the defaults.

grouping: Dimension | None = None class-attribute instance-attribute

What splits each bar into a group of bars.

One of the groupers - ModelDimension, AlgorithmDimension, FeatureDimension, ParameterDimension - or None, which leaves the bars ungrouped.

metric_cls: MetricClass property

The metric this plot reads, taken from its @plot(...) declaration.

Returns:

  • MetricClass –

    The first declared metric class.

Raises:

  • PlotMetricUndeclaredError –

    If the plot declares no metric, so there is nothing to read.

run(benchmark_results: BenchmarkResultContainer, save_dir: str | None = None) -> None

Generate plot output from benchmark results.

Parameters:

  • benchmark_results (BenchmarkResultContainer) –

    Aggregated benchmark data consumed by the plot implementation.

  • save_dir (str | None, default: None ) –

    Directory to save the figure into, by default None.

draw_into(axes: Axes, benchmark_results: BenchmarkResultContainer) -> None

Run this plot onto an existing axes rather than into a figure of its own.

Everything the plot draws goes through the pyplot state, so making that axes the current one is enough to redirect it. The figure, the files, and the window stay the caller's business - which is what lets several plots share one figure, e.g. the summary grid.

Parameters:

  • axes (Axes) –

    The axes to draw on.

  • benchmark_results (BenchmarkResultContainer) –

    Aggregated benchmark data handed to :meth:run.

resolve_missing(df: pd.DataFrame, column: str, *, by: str | None = None, within: str | None = None) -> tuple[pd.DataFrame, dict[tuple[str, str], int]]

Return the drawable rows and how many of them were not, per category.

A metric with nothing to report says so with a None or an infinity - a time to solution of a run that never reached the optimum is the usual one, since the expected time to something that did not happen is unbounded. Neither is a height a bar can have, and leaving them in poisons the aggregate: one infinity turns the mean of an algorithm into an infinity, and a missing value silently shortens it.

What happens to them is :attr:missing, and by default it is nothing: the plot raises rather than quietly showing a mean over fewer models than it claims. Asked to carry on, it either leaves them out or fills them from the values that could be drawn - Missing(policy="max") puts them just past the tallest bar, which is where "worse than everything here" belongs. Either way they are counted and a warning is logged: a bar resting on half its models is a different statement from one resting on all of them, and that is not visible in the bar itself.

Parameters:

  • df (DataFrame) –

    The plotting data.

  • column (str) –

    Column holding the plotted value.

  • by (str | None, default: None ) –

    Column whose categories the missing values are counted per, e.g. the x-axis of a bar plot. Without one they are counted under "".

  • within (str | None, default: None ) –

    Column that splits those categories further, e.g. the grouping of a bar plot. Counting per group is what lets a figure mark the one bar of a group that lost values rather than the whole category it sits in.

Returns:

  • tuple[DataFrame, dict[tuple[str, str], int]] –

    The rows to draw, and the number of missing values per category and group - the group is "" where there is none. The mapping is empty when nothing was missing.

Raises:

  • PlotMissingValuesError –

    If values are missing and the policy is "raise".

place_legend(axes: Axes, handles: list[Any] | None = None, labels: list[str] | None = None) -> None

Put the legend beside the axes, whatever drew it.

Outside the axes for a figure of its own: a legend inside sits on top of the data, and which corner is free depends on the run rather than on the plot - the figure would move its own key around as the numbers change. Beside it, the key is always in the same place and covers nothing.

A panel of someone else's figure is the exception. The room beside it belongs to the panel next to it, so a key anchored there is drawn over a neighbour rather than over the data; inside the panel it stays within the space the plot was given.

Parameters:

  • axes (Axes) –

    The axes the plot was drawn on.

  • handles (list[Any] | None, default: None ) –

    Legend handles, by default the ones already on the axes.

  • labels (list[str] | None, default: None ) –

    Their labels, by default the ones already on the axes.

note_missing(handles: list[Any], labels: list[str], missing: dict[tuple[str, str], int]) -> None

Add the legend entry that says how many values the figure could not draw.

What a plot can say beyond that depends on what it draws. A bar has a slot of its own to put a cross under, so BarPlot marks the categories themselves; a point in a cloud or a step of a sweep has no slot, and the count in the key is the whole statement there - enough that a filled value is not read as a measured one.

Parameters:

  • handles (list[Any]) –

    Legend handles, extended in place.

  • labels (list[str]) –

    Their labels, extended in place alongside handles.

  • missing (dict[tuple[str, str], int]) –

    Number of missing values per category and group, as counted by :meth:resolve_missing.

apply_theme() -> None

Install the seaborn theme this plot is drawn under, unless it has none.

The theme is matplotlib's global state rather than a property of one figure, so it is installed before the figure is built and left in place afterwards: a benchmark themes its plots by handing every one of them the same Theme, not by each plot putting the previous look back.

apply_grid(axes: Axes) -> None

Draw the gridlines the theme asks for, behind everything else on axes.

Parameters:

  • axes (Axes) –

    The axes the plot was drawn on.

setup_figure() -> None

Create a matplotlib figure, unless the plot is drawing into a shared axes.

save_figure(save_dir: str) -> list[Path]

Write the current figure to save_dir once per configured file format.

Parameters:

  • save_dir (str) –

    Directory to save the figure into. Created if it does not exist.

Returns:

  • list[Path] –

    Paths that were written successfully.

finalize_plot(xlabel: str, ylabel: str, title: str, ylim: tuple[float, float] | None = None, x_rotation: int = 45, save_dir: str | None = None) -> None

Apply common axis labels, title, limits, and display behavior.

Parameters:

  • xlabel (str) –

    Label for the x-axis.

  • ylabel (str) –

    Label for the y-axis.

  • title (str) –

    Plot title.

  • ylim (tuple[float, float] | None, default: None ) –

    Lower and upper y-axis limits, by default None.

  • x_rotation (int, default: 45 ) –

    Rotation angle for x-axis tick labels, by default 45.

  • save_dir (str | None, default: None ) –

    Directory to save the figure into, by default None.

apply_grouping(benchmark_results: BenchmarkResultContainer, rows: list[dict[str, Any]]) -> dict[str, Any]

Split rows into groups along :attr:grouping.

What that means is the grouper's business - a column of the plotted data, a value looked up per model, or a setting the algorithms were configured with - and so is deciding that it does not apply, in which case the bars stay ungrouped.

Parameters:

  • benchmark_results (BenchmarkResultContainer) –

    Benchmark data the feature results and algorithm configurations are read from.

  • rows (list[dict[str, Any]]) –

    Row-oriented plot data, annotated - and, where a grouping applies to only part of the data, reduced - in place.

Returns:

  • dict[str, Any] –

    Keyword arguments to forward to :meth:create. Empty when no grouping applies, so call sites can splat it unconditionally.

draw(*, benchmark_results: BenchmarkResultContainer, rows: list[dict[str, Any]], save_dir: str | None = None, **overrides: Any) -> None

Group rows and draw them with the display configuration of this plot.

This is what turns the declared fields - :attr:x, :attr:title, :attr:hline and the rest - into a :meth:create call, so a subclass only has to say which rows it plots. Doing it in one place is also what keeps :attr:group_by working for every bar plot rather than for those that remember to apply it.

Parameters:

  • benchmark_results (BenchmarkResultContainer) –

    Benchmark data, used to look up the groups of a feature :attr:group_by.

  • rows (list[dict[str, Any]]) –

    Row-oriented plot data.

  • save_dir (str | None, default: None ) –

    Directory to save the figure into, by default None.

  • **overrides (Any, default: {} ) –

    Keyword arguments forwarded to :meth:create, overriding the fields.

transform_rows(rows: list[dict[str, Any]], x: str | None, group: str | None) -> list[dict[str, Any]]

Return the rows to plot, by default the rows as they are.

A subclass that has to reduce its rows before they are drawn - pooling counts into a single ratio, say - overrides this rather than :meth:run, so it keeps the shared grouping and display handling. It is told what the bars and the groups turned out to be, since that is what a row has to keep to stay one of them.

Parameters:

  • rows (list[dict[str, Any]]) –

    Row-oriented plot data, already annotated with the dimensions' columns.

  • x (str | None) –

    Column the bars are drawn per, or None when the plot has no rows.

  • group (str | None) –

    Column the bars are split by, or None when they are ungrouped.

Returns:

create(*, rows: list[dict[str, Any]], xlabel: str, ylabel: str, title: str, x: str = 'x', y: str = 'y', aggregation: Aggregation = Aggregation.MEAN, errorbar: ErrorBar | str = AUTO_ERRORBAR, hue: str | None = None, hline: float | None = None, hline_label: str | None = None, hcolor: str = REFERENCE_LINE_COLOUR, baseline: float | None = None, ylim: tuple[float, float] | None = None, legend: bool = False, save_dir: str | None = None, **kwargs: Any) -> None

Create a bar plot from row-oriented data.

Parameters:

  • rows (dict[str, Any]) –

    Row-oriented mapping used to construct the plotting DataFrame.

  • xlabel (str) –

    Label for the x-axis.

  • ylabel (str) –

    Label for the y-axis.

  • title (str) –

    Plot title.

  • x (str, default: 'x' ) –

    Column name mapped to the x-axis, by default "x".

  • y (str, default: 'y' ) –

    Column name mapped to the y-axis, by default "y".

  • aggregation (Aggregation, default: MEAN ) –

    Aggregation strategy applied by seaborn, by default Aggregation.MEAN.

  • errorbar (ErrorBar | str, default: AUTO_ERRORBAR ) –

    Seaborn error bar specification ("sd", ("ci", 95), None to disable). By default "auto", which takes the error bar from aggregation: the spread of the samples for means, none for min/max.

  • hue (str | None, default: None ) –

    Optional grouping column for grouped bars, by default None.

  • hline (float | None, default: None ) –

    Optional horizontal reference line value, by default None.

  • hline_label (str | None, default: None ) –

    Legend label for the horizontal reference line, by default None.

  • hcolor (str, default: REFERENCE_LINE_COLOUR ) –

    Colour of the horizontal reference line, by default black.

  • baseline (float | None, default: None ) –

    Height of a solid black baseline marking where the bars start, by default None. Unlike hline it carries no label and stays out of the legend - it says where zero is, it does not name a target.

  • ylim (tuple[float, float] | None, default: None ) –

    Lower and upper y-axis limits, by default None.

  • legend (bool, default: False ) –

    Whether seaborn should create a legend for hue groups, by default False.

  • save_dir (str | None, default: None ) –

    Directory to save the figure into, by default None.

  • **kwargs (Any, default: {} ) –

    Additional keyword arguments forwarded to :func:seaborn.barplot. They override the defaults computed here, so anything seaborn understands (palette, saturation, capsize, err_kws, ...) can be tuned from the call site.

annotation_text(value: float) -> str

Return the text written above a bar of value.

Applies :attr:annotate_format, unless :attr:annotate_max_decimals allows the value to be written as a plain decimal instead of in scientific notation. A value read off a percent axis is written as a percentage, so the annotation says the same thing as the axis it stands on - unless a format was asked for, which wins.

Parameters:

  • value (float) –

    The aggregated value of one bar.

Returns:

  • str –

    The annotation, e.g. "0.000057" rather than "5.67e-05".

value(metric_result: MetricResult) -> float

Return the number a single metric result contributes.

Parameters:

  • metric_result (MetricResult) –

    One result of :attr:metric_cls.

Returns:

  • float –

    The value plotted for this result, by default the attribute :attr:y names.

rows(benchmark_results: BenchmarkResultContainer) -> list[dict[str, Any]]

Return one row per model and algorithm.

Parameters:

Returns:

  • list[dict[str, Any]] –

    Rows carrying the algorithm, the model, and the plotted value under :attr:y.

FeasibleSampleRatioPlot

Bases: MetricBarPlot

Bar chart of the feasible samples an algorithm produced over all its samples.

The counts are pooled before dividing: every sample of every model enters the same sum, so an algorithm that returns a thousand samples on one model and ten on another is judged by its samples rather than by its models. That is the difference to FeasibilityRatioPlot, which averages the per-model ratios and therefore weights both models equally. 100% - marked by the reference line - means every sample was feasible.

Because each bar is a single pooled number there is no spread to show, so the bars carry no error bars.

Requires the FeasibleSamples metric.

Every display option is inherited and can be set when the plot is constructed, e.g. FeasibleSampleRatioPlot(annotate=False, file_formats=("pgf", "png")). BarPlot documents the colours, value annotations and grouping by a feature; SeabornPlot the figure size and the output formats.

Attributes:

  • figure_filename (str) –

    Stem of the written figure files, by default "feasible_sample_ratio".

Examples:

>>> bench.add_metric(name="feasible_samples", metric=FeasibleSamples())
>>> bench.add_plot(name="feasible_share", plot=FeasibleSampleRatioPlot())
See Also

FeasibilityRatioPlot : The same share averaged per model instead of pooled. FeasibleSolutionFoundPlot : Whether any feasible sample was found, per model.

errorbars: ErrorBars | None = None class-attribute instance-attribute

No error bar: each bar is a single pooled number, so there is no spread to show.

theme: Theme | None = Theme() class-attribute instance-attribute

The seaborn theme the figure is drawn under, and the gridlines behind the marks.

A figure is read against its axis, so it is drawn with the lines that make that possible unless asked otherwise. None takes the theme and the grid away.

missing: Missing = Missing() class-attribute instance-attribute

What becomes of the values the plot cannot draw, and how it says they were there.

option_bundles: dict[str, type[OptionBundle]] = {'style': PlotStyle} class-attribute

Constructor arguments that configure several fields at once, and the bundle each takes.

A style is spread over the bundle fields rather than stored, so a benchmark can hand the same look to every plot while each keeps what it says itself. Applied least specific first: the shared style, then a bundle passed to the plot, then a flat option.

x: Dimension = AlgorithmDimension() class-attribute instance-attribute

What the bars are: one per value of this dimension, and its title on the axis.

aggregation: Aggregation = Aggregation.MEAN class-attribute instance-attribute

Aggregation applied to the values of an x category, by default their mean.

annotation: Annotation | None = None class-attribute instance-attribute

The values written above the bars - how they are formatted and how large.

None, the default, writes none: a bar chart is read off its axis, and a number above every bar is worth its clutter only when the exact value is the point. Pass an Annotation to turn them on, empty for the defaults.

grouping: Dimension | None = None class-attribute instance-attribute

What splits each bar into a group of bars.

One of the groupers - ModelDimension, AlgorithmDimension, FeatureDimension, ParameterDimension - or None, which leaves the bars ungrouped.

metric_cls: MetricClass property

The metric this plot reads, taken from its @plot(...) declaration.

Returns:

  • MetricClass –

    The first declared metric class.

Raises:

  • PlotMetricUndeclaredError –

    If the plot declares no metric, so there is nothing to read.

rows(benchmark_results: BenchmarkResultContainer) -> list[dict[str, Any]]

Return one row per model and algorithm, carrying both sample counts.

Parameters:

Returns:

  • list[dict[str, Any]] –

    Rows to be pooled into one bar per algorithm by :meth:transform_rows.

transform_rows(rows: list[dict[str, Any]], x: str | None, group: str | None) -> list[dict[str, Any]]

Pool the per-model counts into one ratio per bar.

Parameters:

  • rows (list[dict[str, Any]]) –

    One row per model and algorithm, carrying both sample counts.

  • x (str | None) –

    Column the bars are drawn per.

  • group (str | None) –

    Column the bars are split by, or None when they are ungrouped.

Returns:

  • list[dict[str, Any]] –

    One row per bar, in the order the bars first appear.

run(benchmark_results: BenchmarkResultContainer, save_dir: str | None = None) -> None

Generate plot output from benchmark results.

Parameters:

  • benchmark_results (BenchmarkResultContainer) –

    Aggregated benchmark data consumed by the plot implementation.

  • save_dir (str | None, default: None ) –

    Directory to save the figure into, by default None.

draw_into(axes: Axes, benchmark_results: BenchmarkResultContainer) -> None

Run this plot onto an existing axes rather than into a figure of its own.

Everything the plot draws goes through the pyplot state, so making that axes the current one is enough to redirect it. The figure, the files, and the window stay the caller's business - which is what lets several plots share one figure, e.g. the summary grid.

Parameters:

  • axes (Axes) –

    The axes to draw on.

  • benchmark_results (BenchmarkResultContainer) –

    Aggregated benchmark data handed to :meth:run.

resolve_missing(df: pd.DataFrame, column: str, *, by: str | None = None, within: str | None = None) -> tuple[pd.DataFrame, dict[tuple[str, str], int]]

Return the drawable rows and how many of them were not, per category.

A metric with nothing to report says so with a None or an infinity - a time to solution of a run that never reached the optimum is the usual one, since the expected time to something that did not happen is unbounded. Neither is a height a bar can have, and leaving them in poisons the aggregate: one infinity turns the mean of an algorithm into an infinity, and a missing value silently shortens it.

What happens to them is :attr:missing, and by default it is nothing: the plot raises rather than quietly showing a mean over fewer models than it claims. Asked to carry on, it either leaves them out or fills them from the values that could be drawn - Missing(policy="max") puts them just past the tallest bar, which is where "worse than everything here" belongs. Either way they are counted and a warning is logged: a bar resting on half its models is a different statement from one resting on all of them, and that is not visible in the bar itself.

Parameters:

  • df (DataFrame) –

    The plotting data.

  • column (str) –

    Column holding the plotted value.

  • by (str | None, default: None ) –

    Column whose categories the missing values are counted per, e.g. the x-axis of a bar plot. Without one they are counted under "".

  • within (str | None, default: None ) –

    Column that splits those categories further, e.g. the grouping of a bar plot. Counting per group is what lets a figure mark the one bar of a group that lost values rather than the whole category it sits in.

Returns:

  • tuple[DataFrame, dict[tuple[str, str], int]] –

    The rows to draw, and the number of missing values per category and group - the group is "" where there is none. The mapping is empty when nothing was missing.

Raises:

  • PlotMissingValuesError –

    If values are missing and the policy is "raise".

place_legend(axes: Axes, handles: list[Any] | None = None, labels: list[str] | None = None) -> None

Put the legend beside the axes, whatever drew it.

Outside the axes for a figure of its own: a legend inside sits on top of the data, and which corner is free depends on the run rather than on the plot - the figure would move its own key around as the numbers change. Beside it, the key is always in the same place and covers nothing.

A panel of someone else's figure is the exception. The room beside it belongs to the panel next to it, so a key anchored there is drawn over a neighbour rather than over the data; inside the panel it stays within the space the plot was given.

Parameters:

  • axes (Axes) –

    The axes the plot was drawn on.

  • handles (list[Any] | None, default: None ) –

    Legend handles, by default the ones already on the axes.

  • labels (list[str] | None, default: None ) –

    Their labels, by default the ones already on the axes.

note_missing(handles: list[Any], labels: list[str], missing: dict[tuple[str, str], int]) -> None

Add the legend entry that says how many values the figure could not draw.

What a plot can say beyond that depends on what it draws. A bar has a slot of its own to put a cross under, so BarPlot marks the categories themselves; a point in a cloud or a step of a sweep has no slot, and the count in the key is the whole statement there - enough that a filled value is not read as a measured one.

Parameters:

  • handles (list[Any]) –

    Legend handles, extended in place.

  • labels (list[str]) –

    Their labels, extended in place alongside handles.

  • missing (dict[tuple[str, str], int]) –

    Number of missing values per category and group, as counted by :meth:resolve_missing.

apply_theme() -> None

Install the seaborn theme this plot is drawn under, unless it has none.

The theme is matplotlib's global state rather than a property of one figure, so it is installed before the figure is built and left in place afterwards: a benchmark themes its plots by handing every one of them the same Theme, not by each plot putting the previous look back.

apply_grid(axes: Axes) -> None

Draw the gridlines the theme asks for, behind everything else on axes.

Parameters:

  • axes (Axes) –

    The axes the plot was drawn on.

setup_figure() -> None

Create a matplotlib figure, unless the plot is drawing into a shared axes.

save_figure(save_dir: str) -> list[Path]

Write the current figure to save_dir once per configured file format.

Parameters:

  • save_dir (str) –

    Directory to save the figure into. Created if it does not exist.

Returns:

  • list[Path] –

    Paths that were written successfully.

finalize_plot(xlabel: str, ylabel: str, title: str, ylim: tuple[float, float] | None = None, x_rotation: int = 45, save_dir: str | None = None) -> None

Apply common axis labels, title, limits, and display behavior.

Parameters:

  • xlabel (str) –

    Label for the x-axis.

  • ylabel (str) –

    Label for the y-axis.

  • title (str) –

    Plot title.

  • ylim (tuple[float, float] | None, default: None ) –

    Lower and upper y-axis limits, by default None.

  • x_rotation (int, default: 45 ) –

    Rotation angle for x-axis tick labels, by default 45.

  • save_dir (str | None, default: None ) –

    Directory to save the figure into, by default None.

apply_grouping(benchmark_results: BenchmarkResultContainer, rows: list[dict[str, Any]]) -> dict[str, Any]

Split rows into groups along :attr:grouping.

What that means is the grouper's business - a column of the plotted data, a value looked up per model, or a setting the algorithms were configured with - and so is deciding that it does not apply, in which case the bars stay ungrouped.

Parameters:

  • benchmark_results (BenchmarkResultContainer) –

    Benchmark data the feature results and algorithm configurations are read from.

  • rows (list[dict[str, Any]]) –

    Row-oriented plot data, annotated - and, where a grouping applies to only part of the data, reduced - in place.

Returns:

  • dict[str, Any] –

    Keyword arguments to forward to :meth:create. Empty when no grouping applies, so call sites can splat it unconditionally.

draw(*, benchmark_results: BenchmarkResultContainer, rows: list[dict[str, Any]], save_dir: str | None = None, **overrides: Any) -> None

Group rows and draw them with the display configuration of this plot.

This is what turns the declared fields - :attr:x, :attr:title, :attr:hline and the rest - into a :meth:create call, so a subclass only has to say which rows it plots. Doing it in one place is also what keeps :attr:group_by working for every bar plot rather than for those that remember to apply it.

Parameters:

  • benchmark_results (BenchmarkResultContainer) –

    Benchmark data, used to look up the groups of a feature :attr:group_by.

  • rows (list[dict[str, Any]]) –

    Row-oriented plot data.

  • save_dir (str | None, default: None ) –

    Directory to save the figure into, by default None.

  • **overrides (Any, default: {} ) –

    Keyword arguments forwarded to :meth:create, overriding the fields.

create(*, rows: list[dict[str, Any]], xlabel: str, ylabel: str, title: str, x: str = 'x', y: str = 'y', aggregation: Aggregation = Aggregation.MEAN, errorbar: ErrorBar | str = AUTO_ERRORBAR, hue: str | None = None, hline: float | None = None, hline_label: str | None = None, hcolor: str = REFERENCE_LINE_COLOUR, baseline: float | None = None, ylim: tuple[float, float] | None = None, legend: bool = False, save_dir: str | None = None, **kwargs: Any) -> None

Create a bar plot from row-oriented data.

Parameters:

  • rows (dict[str, Any]) –

    Row-oriented mapping used to construct the plotting DataFrame.

  • xlabel (str) –

    Label for the x-axis.

  • ylabel (str) –

    Label for the y-axis.

  • title (str) –

    Plot title.

  • x (str, default: 'x' ) –

    Column name mapped to the x-axis, by default "x".

  • y (str, default: 'y' ) –

    Column name mapped to the y-axis, by default "y".

  • aggregation (Aggregation, default: MEAN ) –

    Aggregation strategy applied by seaborn, by default Aggregation.MEAN.

  • errorbar (ErrorBar | str, default: AUTO_ERRORBAR ) –

    Seaborn error bar specification ("sd", ("ci", 95), None to disable). By default "auto", which takes the error bar from aggregation: the spread of the samples for means, none for min/max.

  • hue (str | None, default: None ) –

    Optional grouping column for grouped bars, by default None.

  • hline (float | None, default: None ) –

    Optional horizontal reference line value, by default None.

  • hline_label (str | None, default: None ) –

    Legend label for the horizontal reference line, by default None.

  • hcolor (str, default: REFERENCE_LINE_COLOUR ) –

    Colour of the horizontal reference line, by default black.

  • baseline (float | None, default: None ) –

    Height of a solid black baseline marking where the bars start, by default None. Unlike hline it carries no label and stays out of the legend - it says where zero is, it does not name a target.

  • ylim (tuple[float, float] | None, default: None ) –

    Lower and upper y-axis limits, by default None.

  • legend (bool, default: False ) –

    Whether seaborn should create a legend for hue groups, by default False.

  • save_dir (str | None, default: None ) –

    Directory to save the figure into, by default None.

  • **kwargs (Any, default: {} ) –

    Additional keyword arguments forwarded to :func:seaborn.barplot. They override the defaults computed here, so anything seaborn understands (palette, saturation, capsize, err_kws, ...) can be tuned from the call site.

annotation_text(value: float) -> str

Return the text written above a bar of value.

Applies :attr:annotate_format, unless :attr:annotate_max_decimals allows the value to be written as a plain decimal instead of in scientific notation. A value read off a percent axis is written as a percentage, so the annotation says the same thing as the axis it stands on - unless a format was asked for, which wins.

Parameters:

  • value (float) –

    The aggregated value of one bar.

Returns:

  • str –

    The annotation, e.g. "0.000057" rather than "5.67e-05".

value(metric_result: MetricResult) -> float

Return the number a single metric result contributes.

Parameters:

  • metric_result (MetricResult) –

    One result of :attr:metric_cls.

Returns:

  • float –

    The value plotted for this result, by default the attribute :attr:y names.

FeasibleSolutionFoundPlot

Bases: MetricBarPlot

Bar chart showing the percentage of models solved feasibly per algorithm.

A model counts as solved when at least one sample was feasible, i.e. when the FeasibilityRatio is greater than zero. Averaging that indicator over all models gives the share of the benchmark an algorithm could handle at all, independent of how many of its samples were feasible.

Every display option is inherited and can be set when the plot is constructed, e.g. FeasibleSolutionFoundPlot(annotate=False, file_formats=("pgf", "png")). BarPlot documents the colours, error bars, value annotations and grouping by a feature; SeabornPlot the figure size and the output formats.

Attributes:

  • figure_filename (str) –

    Stem of the written figure files, by default "feasible_solution_found".

  • annotation (Annotation) –

    The values written above the bars, in percent by default.

Examples:

>>> bench.add_metric(name="feasibility", metric=FeasibilityRatio())
>>> bench.add_plot(name="feasible_found", plot=FeasibleSolutionFoundPlot())

theme: Theme | None = Theme() class-attribute instance-attribute

The seaborn theme the figure is drawn under, and the gridlines behind the marks.

A figure is read against its axis, so it is drawn with the lines that make that possible unless asked otherwise. None takes the theme and the grid away.

missing: Missing = Missing() class-attribute instance-attribute

What becomes of the values the plot cannot draw, and how it says they were there.

option_bundles: dict[str, type[OptionBundle]] = {'style': PlotStyle} class-attribute

Constructor arguments that configure several fields at once, and the bundle each takes.

A style is spread over the bundle fields rather than stored, so a benchmark can hand the same look to every plot while each keeps what it says itself. Applied least specific first: the shared style, then a bundle passed to the plot, then a flat option.

x: Dimension = AlgorithmDimension() class-attribute instance-attribute

What the bars are: one per value of this dimension, and its title on the axis.

aggregation: Aggregation = Aggregation.MEAN class-attribute instance-attribute

Aggregation applied to the values of an x category, by default their mean.

errorbars: ErrorBars | None = ErrorBars() class-attribute instance-attribute

The error bars drawn on top of the bars: what they show, their colour and caps.

None draws none, the same as ErrorBars(spec=None).

grouping: Dimension | None = None class-attribute instance-attribute

What splits each bar into a group of bars.

One of the groupers - ModelDimension, AlgorithmDimension, FeatureDimension, ParameterDimension - or None, which leaves the bars ungrouped.

metric_cls: MetricClass property

The metric this plot reads, taken from its @plot(...) declaration.

Returns:

  • MetricClass –

    The first declared metric class.

Raises:

  • PlotMetricUndeclaredError –

    If the plot declares no metric, so there is nothing to read.

value(metric_result: MetricResult) -> float

Return 100 when the algorithm found a feasible sample for this model, 0 otherwise.

A result that reports no ratio at all is neither: "the metric has nothing to say about this model" is not the same statement as "this algorithm found nothing feasible", and turning the first into the second would count a gap in the data as a failure of the solver. It stays the missing value it is, for Missing to decide about with the rest of them.

Parameters:

  • metric_result (MetricResult) –

    The feasibility ratio of one model and algorithm.

Returns:

  • float –

    The indicator, in percent, so the bars average into the share of the benchmark an algorithm could handle at all, or nan where there is none.

run(benchmark_results: BenchmarkResultContainer, save_dir: str | None = None) -> None

Generate plot output from benchmark results.

Parameters:

  • benchmark_results (BenchmarkResultContainer) –

    Aggregated benchmark data consumed by the plot implementation.

  • save_dir (str | None, default: None ) –

    Directory to save the figure into, by default None.

draw_into(axes: Axes, benchmark_results: BenchmarkResultContainer) -> None

Run this plot onto an existing axes rather than into a figure of its own.

Everything the plot draws goes through the pyplot state, so making that axes the current one is enough to redirect it. The figure, the files, and the window stay the caller's business - which is what lets several plots share one figure, e.g. the summary grid.

Parameters:

  • axes (Axes) –

    The axes to draw on.

  • benchmark_results (BenchmarkResultContainer) –

    Aggregated benchmark data handed to :meth:run.

resolve_missing(df: pd.DataFrame, column: str, *, by: str | None = None, within: str | None = None) -> tuple[pd.DataFrame, dict[tuple[str, str], int]]

Return the drawable rows and how many of them were not, per category.

A metric with nothing to report says so with a None or an infinity - a time to solution of a run that never reached the optimum is the usual one, since the expected time to something that did not happen is unbounded. Neither is a height a bar can have, and leaving them in poisons the aggregate: one infinity turns the mean of an algorithm into an infinity, and a missing value silently shortens it.

What happens to them is :attr:missing, and by default it is nothing: the plot raises rather than quietly showing a mean over fewer models than it claims. Asked to carry on, it either leaves them out or fills them from the values that could be drawn - Missing(policy="max") puts them just past the tallest bar, which is where "worse than everything here" belongs. Either way they are counted and a warning is logged: a bar resting on half its models is a different statement from one resting on all of them, and that is not visible in the bar itself.

Parameters:

  • df (DataFrame) –

    The plotting data.

  • column (str) –

    Column holding the plotted value.

  • by (str | None, default: None ) –

    Column whose categories the missing values are counted per, e.g. the x-axis of a bar plot. Without one they are counted under "".

  • within (str | None, default: None ) –

    Column that splits those categories further, e.g. the grouping of a bar plot. Counting per group is what lets a figure mark the one bar of a group that lost values rather than the whole category it sits in.

Returns:

  • tuple[DataFrame, dict[tuple[str, str], int]] –

    The rows to draw, and the number of missing values per category and group - the group is "" where there is none. The mapping is empty when nothing was missing.

Raises:

  • PlotMissingValuesError –

    If values are missing and the policy is "raise".

place_legend(axes: Axes, handles: list[Any] | None = None, labels: list[str] | None = None) -> None

Put the legend beside the axes, whatever drew it.

Outside the axes for a figure of its own: a legend inside sits on top of the data, and which corner is free depends on the run rather than on the plot - the figure would move its own key around as the numbers change. Beside it, the key is always in the same place and covers nothing.

A panel of someone else's figure is the exception. The room beside it belongs to the panel next to it, so a key anchored there is drawn over a neighbour rather than over the data; inside the panel it stays within the space the plot was given.

Parameters:

  • axes (Axes) –

    The axes the plot was drawn on.

  • handles (list[Any] | None, default: None ) –

    Legend handles, by default the ones already on the axes.

  • labels (list[str] | None, default: None ) –

    Their labels, by default the ones already on the axes.

note_missing(handles: list[Any], labels: list[str], missing: dict[tuple[str, str], int]) -> None

Add the legend entry that says how many values the figure could not draw.

What a plot can say beyond that depends on what it draws. A bar has a slot of its own to put a cross under, so BarPlot marks the categories themselves; a point in a cloud or a step of a sweep has no slot, and the count in the key is the whole statement there - enough that a filled value is not read as a measured one.

Parameters:

  • handles (list[Any]) –

    Legend handles, extended in place.

  • labels (list[str]) –

    Their labels, extended in place alongside handles.

  • missing (dict[tuple[str, str], int]) –

    Number of missing values per category and group, as counted by :meth:resolve_missing.

apply_theme() -> None

Install the seaborn theme this plot is drawn under, unless it has none.

The theme is matplotlib's global state rather than a property of one figure, so it is installed before the figure is built and left in place afterwards: a benchmark themes its plots by handing every one of them the same Theme, not by each plot putting the previous look back.

apply_grid(axes: Axes) -> None

Draw the gridlines the theme asks for, behind everything else on axes.

Parameters:

  • axes (Axes) –

    The axes the plot was drawn on.

setup_figure() -> None

Create a matplotlib figure, unless the plot is drawing into a shared axes.

save_figure(save_dir: str) -> list[Path]

Write the current figure to save_dir once per configured file format.

Parameters:

  • save_dir (str) –

    Directory to save the figure into. Created if it does not exist.

Returns:

  • list[Path] –

    Paths that were written successfully.

finalize_plot(xlabel: str, ylabel: str, title: str, ylim: tuple[float, float] | None = None, x_rotation: int = 45, save_dir: str | None = None) -> None

Apply common axis labels, title, limits, and display behavior.

Parameters:

  • xlabel (str) –

    Label for the x-axis.

  • ylabel (str) –

    Label for the y-axis.

  • title (str) –

    Plot title.

  • ylim (tuple[float, float] | None, default: None ) –

    Lower and upper y-axis limits, by default None.

  • x_rotation (int, default: 45 ) –

    Rotation angle for x-axis tick labels, by default 45.

  • save_dir (str | None, default: None ) –

    Directory to save the figure into, by default None.

apply_grouping(benchmark_results: BenchmarkResultContainer, rows: list[dict[str, Any]]) -> dict[str, Any]

Split rows into groups along :attr:grouping.

What that means is the grouper's business - a column of the plotted data, a value looked up per model, or a setting the algorithms were configured with - and so is deciding that it does not apply, in which case the bars stay ungrouped.

Parameters:

  • benchmark_results (BenchmarkResultContainer) –

    Benchmark data the feature results and algorithm configurations are read from.

  • rows (list[dict[str, Any]]) –

    Row-oriented plot data, annotated - and, where a grouping applies to only part of the data, reduced - in place.

Returns:

  • dict[str, Any] –

    Keyword arguments to forward to :meth:create. Empty when no grouping applies, so call sites can splat it unconditionally.

draw(*, benchmark_results: BenchmarkResultContainer, rows: list[dict[str, Any]], save_dir: str | None = None, **overrides: Any) -> None

Group rows and draw them with the display configuration of this plot.

This is what turns the declared fields - :attr:x, :attr:title, :attr:hline and the rest - into a :meth:create call, so a subclass only has to say which rows it plots. Doing it in one place is also what keeps :attr:group_by working for every bar plot rather than for those that remember to apply it.

Parameters:

  • benchmark_results (BenchmarkResultContainer) –

    Benchmark data, used to look up the groups of a feature :attr:group_by.

  • rows (list[dict[str, Any]]) –

    Row-oriented plot data.

  • save_dir (str | None, default: None ) –

    Directory to save the figure into, by default None.

  • **overrides (Any, default: {} ) –

    Keyword arguments forwarded to :meth:create, overriding the fields.

transform_rows(rows: list[dict[str, Any]], x: str | None, group: str | None) -> list[dict[str, Any]]

Return the rows to plot, by default the rows as they are.

A subclass that has to reduce its rows before they are drawn - pooling counts into a single ratio, say - overrides this rather than :meth:run, so it keeps the shared grouping and display handling. It is told what the bars and the groups turned out to be, since that is what a row has to keep to stay one of them.

Parameters:

  • rows (list[dict[str, Any]]) –

    Row-oriented plot data, already annotated with the dimensions' columns.

  • x (str | None) –

    Column the bars are drawn per, or None when the plot has no rows.

  • group (str | None) –

    Column the bars are split by, or None when they are ungrouped.

Returns:

create(*, rows: list[dict[str, Any]], xlabel: str, ylabel: str, title: str, x: str = 'x', y: str = 'y', aggregation: Aggregation = Aggregation.MEAN, errorbar: ErrorBar | str = AUTO_ERRORBAR, hue: str | None = None, hline: float | None = None, hline_label: str | None = None, hcolor: str = REFERENCE_LINE_COLOUR, baseline: float | None = None, ylim: tuple[float, float] | None = None, legend: bool = False, save_dir: str | None = None, **kwargs: Any) -> None

Create a bar plot from row-oriented data.

Parameters:

  • rows (dict[str, Any]) –

    Row-oriented mapping used to construct the plotting DataFrame.

  • xlabel (str) –

    Label for the x-axis.

  • ylabel (str) –

    Label for the y-axis.

  • title (str) –

    Plot title.

  • x (str, default: 'x' ) –

    Column name mapped to the x-axis, by default "x".

  • y (str, default: 'y' ) –

    Column name mapped to the y-axis, by default "y".

  • aggregation (Aggregation, default: MEAN ) –

    Aggregation strategy applied by seaborn, by default Aggregation.MEAN.

  • errorbar (ErrorBar | str, default: AUTO_ERRORBAR ) –

    Seaborn error bar specification ("sd", ("ci", 95), None to disable). By default "auto", which takes the error bar from aggregation: the spread of the samples for means, none for min/max.

  • hue (str | None, default: None ) –

    Optional grouping column for grouped bars, by default None.

  • hline (float | None, default: None ) –

    Optional horizontal reference line value, by default None.

  • hline_label (str | None, default: None ) –

    Legend label for the horizontal reference line, by default None.

  • hcolor (str, default: REFERENCE_LINE_COLOUR ) –

    Colour of the horizontal reference line, by default black.

  • baseline (float | None, default: None ) –

    Height of a solid black baseline marking where the bars start, by default None. Unlike hline it carries no label and stays out of the legend - it says where zero is, it does not name a target.

  • ylim (tuple[float, float] | None, default: None ) –

    Lower and upper y-axis limits, by default None.

  • legend (bool, default: False ) –

    Whether seaborn should create a legend for hue groups, by default False.

  • save_dir (str | None, default: None ) –

    Directory to save the figure into, by default None.

  • **kwargs (Any, default: {} ) –

    Additional keyword arguments forwarded to :func:seaborn.barplot. They override the defaults computed here, so anything seaborn understands (palette, saturation, capsize, err_kws, ...) can be tuned from the call site.

annotation_text(value: float) -> str

Return the text written above a bar of value.

Applies :attr:annotate_format, unless :attr:annotate_max_decimals allows the value to be written as a plain decimal instead of in scientific notation. A value read off a percent axis is written as a percentage, so the annotation says the same thing as the axis it stands on - unless a format was asked for, which wins.

Parameters:

  • value (float) –

    The aggregated value of one bar.

Returns:

  • str –

    The annotation, e.g. "0.000057" rather than "5.67e-05".

rows(benchmark_results: BenchmarkResultContainer) -> list[dict[str, Any]]

Return one row per model and algorithm.

Parameters:

Returns:

  • list[dict[str, Any]] –

    Rows carrying the algorithm, the model, and the plotted value under :attr:y.

FractionOfOverallBestSolutionPlot

Bases: MetricBarPlot

Bar chart of the fraction of the overall best solution per algorithm.

Each algorithm's objective value is measured against the best value any algorithm in the benchmark reached for that model, so 1.0 - marked by the reference line - means it was the best of the field. Useful when no optimum is known. Aggregated over every model - the mean by default - with the spread as the error bar.

Requires the FractionOfOverallBestSolution metric.

Every display option is inherited and can be set when the plot is constructed, e.g. FractionOfOverallBestSolutionPlot(annotate=False, file_formats=("pgf", "png")). BarPlot documents the colours, error bars, value annotations and grouping by a feature; SeabornPlot the figure size and the output formats.

Attributes:

  • figure_filename (str) –

    Stem of the written figure files, by default "fraction_of_overall_best_solution".

Examples:

>>> bench.add_metric(name="best_found", metric=FractionOfOverallBestSolution())
>>> bench.add_plot(name="avg_best", plot=FractionOfOverallBestSolutionPlot())

theme: Theme | None = Theme() class-attribute instance-attribute

The seaborn theme the figure is drawn under, and the gridlines behind the marks.

A figure is read against its axis, so it is drawn with the lines that make that possible unless asked otherwise. None takes the theme and the grid away.

missing: Missing = Missing() class-attribute instance-attribute

What becomes of the values the plot cannot draw, and how it says they were there.

option_bundles: dict[str, type[OptionBundle]] = {'style': PlotStyle} class-attribute

Constructor arguments that configure several fields at once, and the bundle each takes.

A style is spread over the bundle fields rather than stored, so a benchmark can hand the same look to every plot while each keeps what it says itself. Applied least specific first: the shared style, then a bundle passed to the plot, then a flat option.

x: Dimension = AlgorithmDimension() class-attribute instance-attribute

What the bars are: one per value of this dimension, and its title on the axis.

aggregation: Aggregation = Aggregation.MEAN class-attribute instance-attribute

Aggregation applied to the values of an x category, by default their mean.

errorbars: ErrorBars | None = ErrorBars() class-attribute instance-attribute

The error bars drawn on top of the bars: what they show, their colour and caps.

None draws none, the same as ErrorBars(spec=None).

annotation: Annotation | None = None class-attribute instance-attribute

The values written above the bars - how they are formatted and how large.

None, the default, writes none: a bar chart is read off its axis, and a number above every bar is worth its clutter only when the exact value is the point. Pass an Annotation to turn them on, empty for the defaults.

grouping: Dimension | None = None class-attribute instance-attribute

What splits each bar into a group of bars.

One of the groupers - ModelDimension, AlgorithmDimension, FeatureDimension, ParameterDimension - or None, which leaves the bars ungrouped.

metric_cls: MetricClass property

The metric this plot reads, taken from its @plot(...) declaration.

Returns:

  • MetricClass –

    The first declared metric class.

Raises:

  • PlotMetricUndeclaredError –

    If the plot declares no metric, so there is nothing to read.

run(benchmark_results: BenchmarkResultContainer, save_dir: str | None = None) -> None

Generate plot output from benchmark results.

Parameters:

  • benchmark_results (BenchmarkResultContainer) –

    Aggregated benchmark data consumed by the plot implementation.

  • save_dir (str | None, default: None ) –

    Directory to save the figure into, by default None.

draw_into(axes: Axes, benchmark_results: BenchmarkResultContainer) -> None

Run this plot onto an existing axes rather than into a figure of its own.

Everything the plot draws goes through the pyplot state, so making that axes the current one is enough to redirect it. The figure, the files, and the window stay the caller's business - which is what lets several plots share one figure, e.g. the summary grid.

Parameters:

  • axes (Axes) –

    The axes to draw on.

  • benchmark_results (BenchmarkResultContainer) –

    Aggregated benchmark data handed to :meth:run.

resolve_missing(df: pd.DataFrame, column: str, *, by: str | None = None, within: str | None = None) -> tuple[pd.DataFrame, dict[tuple[str, str], int]]

Return the drawable rows and how many of them were not, per category.

A metric with nothing to report says so with a None or an infinity - a time to solution of a run that never reached the optimum is the usual one, since the expected time to something that did not happen is unbounded. Neither is a height a bar can have, and leaving them in poisons the aggregate: one infinity turns the mean of an algorithm into an infinity, and a missing value silently shortens it.

What happens to them is :attr:missing, and by default it is nothing: the plot raises rather than quietly showing a mean over fewer models than it claims. Asked to carry on, it either leaves them out or fills them from the values that could be drawn - Missing(policy="max") puts them just past the tallest bar, which is where "worse than everything here" belongs. Either way they are counted and a warning is logged: a bar resting on half its models is a different statement from one resting on all of them, and that is not visible in the bar itself.

Parameters:

  • df (DataFrame) –

    The plotting data.

  • column (str) –

    Column holding the plotted value.

  • by (str | None, default: None ) –

    Column whose categories the missing values are counted per, e.g. the x-axis of a bar plot. Without one they are counted under "".

  • within (str | None, default: None ) –

    Column that splits those categories further, e.g. the grouping of a bar plot. Counting per group is what lets a figure mark the one bar of a group that lost values rather than the whole category it sits in.

Returns:

  • tuple[DataFrame, dict[tuple[str, str], int]] –

    The rows to draw, and the number of missing values per category and group - the group is "" where there is none. The mapping is empty when nothing was missing.

Raises:

  • PlotMissingValuesError –

    If values are missing and the policy is "raise".

place_legend(axes: Axes, handles: list[Any] | None = None, labels: list[str] | None = None) -> None

Put the legend beside the axes, whatever drew it.

Outside the axes for a figure of its own: a legend inside sits on top of the data, and which corner is free depends on the run rather than on the plot - the figure would move its own key around as the numbers change. Beside it, the key is always in the same place and covers nothing.

A panel of someone else's figure is the exception. The room beside it belongs to the panel next to it, so a key anchored there is drawn over a neighbour rather than over the data; inside the panel it stays within the space the plot was given.

Parameters:

  • axes (Axes) –

    The axes the plot was drawn on.

  • handles (list[Any] | None, default: None ) –

    Legend handles, by default the ones already on the axes.

  • labels (list[str] | None, default: None ) –

    Their labels, by default the ones already on the axes.

note_missing(handles: list[Any], labels: list[str], missing: dict[tuple[str, str], int]) -> None

Add the legend entry that says how many values the figure could not draw.

What a plot can say beyond that depends on what it draws. A bar has a slot of its own to put a cross under, so BarPlot marks the categories themselves; a point in a cloud or a step of a sweep has no slot, and the count in the key is the whole statement there - enough that a filled value is not read as a measured one.

Parameters:

  • handles (list[Any]) –

    Legend handles, extended in place.

  • labels (list[str]) –

    Their labels, extended in place alongside handles.

  • missing (dict[tuple[str, str], int]) –

    Number of missing values per category and group, as counted by :meth:resolve_missing.

apply_theme() -> None

Install the seaborn theme this plot is drawn under, unless it has none.

The theme is matplotlib's global state rather than a property of one figure, so it is installed before the figure is built and left in place afterwards: a benchmark themes its plots by handing every one of them the same Theme, not by each plot putting the previous look back.

apply_grid(axes: Axes) -> None

Draw the gridlines the theme asks for, behind everything else on axes.

Parameters:

  • axes (Axes) –

    The axes the plot was drawn on.

setup_figure() -> None

Create a matplotlib figure, unless the plot is drawing into a shared axes.

save_figure(save_dir: str) -> list[Path]

Write the current figure to save_dir once per configured file format.

Parameters:

  • save_dir (str) –

    Directory to save the figure into. Created if it does not exist.

Returns:

  • list[Path] –

    Paths that were written successfully.

finalize_plot(xlabel: str, ylabel: str, title: str, ylim: tuple[float, float] | None = None, x_rotation: int = 45, save_dir: str | None = None) -> None

Apply common axis labels, title, limits, and display behavior.

Parameters:

  • xlabel (str) –

    Label for the x-axis.

  • ylabel (str) –

    Label for the y-axis.

  • title (str) –

    Plot title.

  • ylim (tuple[float, float] | None, default: None ) –

    Lower and upper y-axis limits, by default None.

  • x_rotation (int, default: 45 ) –

    Rotation angle for x-axis tick labels, by default 45.

  • save_dir (str | None, default: None ) –

    Directory to save the figure into, by default None.

apply_grouping(benchmark_results: BenchmarkResultContainer, rows: list[dict[str, Any]]) -> dict[str, Any]

Split rows into groups along :attr:grouping.

What that means is the grouper's business - a column of the plotted data, a value looked up per model, or a setting the algorithms were configured with - and so is deciding that it does not apply, in which case the bars stay ungrouped.

Parameters:

  • benchmark_results (BenchmarkResultContainer) –

    Benchmark data the feature results and algorithm configurations are read from.

  • rows (list[dict[str, Any]]) –

    Row-oriented plot data, annotated - and, where a grouping applies to only part of the data, reduced - in place.

Returns:

  • dict[str, Any] –

    Keyword arguments to forward to :meth:create. Empty when no grouping applies, so call sites can splat it unconditionally.

draw(*, benchmark_results: BenchmarkResultContainer, rows: list[dict[str, Any]], save_dir: str | None = None, **overrides: Any) -> None

Group rows and draw them with the display configuration of this plot.

This is what turns the declared fields - :attr:x, :attr:title, :attr:hline and the rest - into a :meth:create call, so a subclass only has to say which rows it plots. Doing it in one place is also what keeps :attr:group_by working for every bar plot rather than for those that remember to apply it.

Parameters:

  • benchmark_results (BenchmarkResultContainer) –

    Benchmark data, used to look up the groups of a feature :attr:group_by.

  • rows (list[dict[str, Any]]) –

    Row-oriented plot data.

  • save_dir (str | None, default: None ) –

    Directory to save the figure into, by default None.

  • **overrides (Any, default: {} ) –

    Keyword arguments forwarded to :meth:create, overriding the fields.

transform_rows(rows: list[dict[str, Any]], x: str | None, group: str | None) -> list[dict[str, Any]]

Return the rows to plot, by default the rows as they are.

A subclass that has to reduce its rows before they are drawn - pooling counts into a single ratio, say - overrides this rather than :meth:run, so it keeps the shared grouping and display handling. It is told what the bars and the groups turned out to be, since that is what a row has to keep to stay one of them.

Parameters:

  • rows (list[dict[str, Any]]) –

    Row-oriented plot data, already annotated with the dimensions' columns.

  • x (str | None) –

    Column the bars are drawn per, or None when the plot has no rows.

  • group (str | None) –

    Column the bars are split by, or None when they are ungrouped.

Returns:

create(*, rows: list[dict[str, Any]], xlabel: str, ylabel: str, title: str, x: str = 'x', y: str = 'y', aggregation: Aggregation = Aggregation.MEAN, errorbar: ErrorBar | str = AUTO_ERRORBAR, hue: str | None = None, hline: float | None = None, hline_label: str | None = None, hcolor: str = REFERENCE_LINE_COLOUR, baseline: float | None = None, ylim: tuple[float, float] | None = None, legend: bool = False, save_dir: str | None = None, **kwargs: Any) -> None

Create a bar plot from row-oriented data.

Parameters:

  • rows (dict[str, Any]) –

    Row-oriented mapping used to construct the plotting DataFrame.

  • xlabel (str) –

    Label for the x-axis.

  • ylabel (str) –

    Label for the y-axis.

  • title (str) –

    Plot title.

  • x (str, default: 'x' ) –

    Column name mapped to the x-axis, by default "x".

  • y (str, default: 'y' ) –

    Column name mapped to the y-axis, by default "y".

  • aggregation (Aggregation, default: MEAN ) –

    Aggregation strategy applied by seaborn, by default Aggregation.MEAN.

  • errorbar (ErrorBar | str, default: AUTO_ERRORBAR ) –

    Seaborn error bar specification ("sd", ("ci", 95), None to disable). By default "auto", which takes the error bar from aggregation: the spread of the samples for means, none for min/max.

  • hue (str | None, default: None ) –

    Optional grouping column for grouped bars, by default None.

  • hline (float | None, default: None ) –

    Optional horizontal reference line value, by default None.

  • hline_label (str | None, default: None ) –

    Legend label for the horizontal reference line, by default None.

  • hcolor (str, default: REFERENCE_LINE_COLOUR ) –

    Colour of the horizontal reference line, by default black.

  • baseline (float | None, default: None ) –

    Height of a solid black baseline marking where the bars start, by default None. Unlike hline it carries no label and stays out of the legend - it says where zero is, it does not name a target.

  • ylim (tuple[float, float] | None, default: None ) –

    Lower and upper y-axis limits, by default None.

  • legend (bool, default: False ) –

    Whether seaborn should create a legend for hue groups, by default False.

  • save_dir (str | None, default: None ) –

    Directory to save the figure into, by default None.

  • **kwargs (Any, default: {} ) –

    Additional keyword arguments forwarded to :func:seaborn.barplot. They override the defaults computed here, so anything seaborn understands (palette, saturation, capsize, err_kws, ...) can be tuned from the call site.

annotation_text(value: float) -> str

Return the text written above a bar of value.

Applies :attr:annotate_format, unless :attr:annotate_max_decimals allows the value to be written as a plain decimal instead of in scientific notation. A value read off a percent axis is written as a percentage, so the annotation says the same thing as the axis it stands on - unless a format was asked for, which wins.

Parameters:

  • value (float) –

    The aggregated value of one bar.

Returns:

  • str –

    The annotation, e.g. "0.000057" rather than "5.67e-05".

value(metric_result: MetricResult) -> float

Return the number a single metric result contributes.

Parameters:

  • metric_result (MetricResult) –

    One result of :attr:metric_cls.

Returns:

  • float –

    The value plotted for this result, by default the attribute :attr:y names.

rows(benchmark_results: BenchmarkResultContainer) -> list[dict[str, Any]]

Return one row per model and algorithm.

Parameters:

Returns:

  • list[dict[str, Any]] –

    Rows carrying the algorithm, the model, and the plotted value under :attr:y.

RuntimePerModelPlot

Bases: MetricBarPlot

Bar chart of runtime per model, with one bar per algorithm inside each model.

Keeps the models apart instead of averaging over them, which shows where a single hard instance drives an algorithm's average runtime up.

Requires the Runtime metric.

Every display option is inherited and can be set when the plot is constructed, e.g. RuntimePerModelPlot(annotate=False, file_formats=("pgf", "png")). BarPlot documents the colours, error bars and value annotations; SeabornPlot the figure size and the output formats. group_by does not apply here - this plot already uses the hue channel for the algorithm.

Attributes:

  • figure_filename (str) –

    Stem of the written figure files, by default "runtime_per_model".

Examples:

>>> bench.add_metric(name="runtime", metric=Runtime())
>>> bench.add_plot(name="runtime_per_model", plot=RuntimePerModelPlot())
See Also

RuntimePlot : The same numbers averaged over all models.

y: MetricDimension = MetricDimension('runtime_seconds', 'Runtime (s)', baseline=0.0) class-attribute instance-attribute

Runtime in seconds, standing on a solid line at zero, as on RuntimePlot.

theme: Theme | None = Theme() class-attribute instance-attribute

The seaborn theme the figure is drawn under, and the gridlines behind the marks.

A figure is read against its axis, so it is drawn with the lines that make that possible unless asked otherwise. None takes the theme and the grid away.

missing: Missing = Missing() class-attribute instance-attribute

What becomes of the values the plot cannot draw, and how it says they were there.

option_bundles: dict[str, type[OptionBundle]] = {'style': PlotStyle} class-attribute

Constructor arguments that configure several fields at once, and the bundle each takes.

A style is spread over the bundle fields rather than stored, so a benchmark can hand the same look to every plot while each keeps what it says itself. Applied least specific first: the shared style, then a bundle passed to the plot, then a flat option.

aggregation: Aggregation = Aggregation.MEAN class-attribute instance-attribute

Aggregation applied to the values of an x category, by default their mean.

errorbars: ErrorBars | None = ErrorBars() class-attribute instance-attribute

The error bars drawn on top of the bars: what they show, their colour and caps.

None draws none, the same as ErrorBars(spec=None).

annotation: Annotation | None = None class-attribute instance-attribute

The values written above the bars - how they are formatted and how large.

None, the default, writes none: a bar chart is read off its axis, and a number above every bar is worth its clutter only when the exact value is the point. Pass an Annotation to turn them on, empty for the defaults.

metric_cls: MetricClass property

The metric this plot reads, taken from its @plot(...) declaration.

Returns:

  • MetricClass –

    The first declared metric class.

Raises:

  • PlotMetricUndeclaredError –

    If the plot declares no metric, so there is nothing to read.

run(benchmark_results: BenchmarkResultContainer, save_dir: str | None = None) -> None

Generate plot output from benchmark results.

Parameters:

  • benchmark_results (BenchmarkResultContainer) –

    Aggregated benchmark data consumed by the plot implementation.

  • save_dir (str | None, default: None ) –

    Directory to save the figure into, by default None.

draw_into(axes: Axes, benchmark_results: BenchmarkResultContainer) -> None

Run this plot onto an existing axes rather than into a figure of its own.

Everything the plot draws goes through the pyplot state, so making that axes the current one is enough to redirect it. The figure, the files, and the window stay the caller's business - which is what lets several plots share one figure, e.g. the summary grid.

Parameters:

  • axes (Axes) –

    The axes to draw on.

  • benchmark_results (BenchmarkResultContainer) –

    Aggregated benchmark data handed to :meth:run.

resolve_missing(df: pd.DataFrame, column: str, *, by: str | None = None, within: str | None = None) -> tuple[pd.DataFrame, dict[tuple[str, str], int]]

Return the drawable rows and how many of them were not, per category.

A metric with nothing to report says so with a None or an infinity - a time to solution of a run that never reached the optimum is the usual one, since the expected time to something that did not happen is unbounded. Neither is a height a bar can have, and leaving them in poisons the aggregate: one infinity turns the mean of an algorithm into an infinity, and a missing value silently shortens it.

What happens to them is :attr:missing, and by default it is nothing: the plot raises rather than quietly showing a mean over fewer models than it claims. Asked to carry on, it either leaves them out or fills them from the values that could be drawn - Missing(policy="max") puts them just past the tallest bar, which is where "worse than everything here" belongs. Either way they are counted and a warning is logged: a bar resting on half its models is a different statement from one resting on all of them, and that is not visible in the bar itself.

Parameters:

  • df (DataFrame) –

    The plotting data.

  • column (str) –

    Column holding the plotted value.

  • by (str | None, default: None ) –

    Column whose categories the missing values are counted per, e.g. the x-axis of a bar plot. Without one they are counted under "".

  • within (str | None, default: None ) –

    Column that splits those categories further, e.g. the grouping of a bar plot. Counting per group is what lets a figure mark the one bar of a group that lost values rather than the whole category it sits in.

Returns:

  • tuple[DataFrame, dict[tuple[str, str], int]] –

    The rows to draw, and the number of missing values per category and group - the group is "" where there is none. The mapping is empty when nothing was missing.

Raises:

  • PlotMissingValuesError –

    If values are missing and the policy is "raise".

place_legend(axes: Axes, handles: list[Any] | None = None, labels: list[str] | None = None) -> None

Put the legend beside the axes, whatever drew it.

Outside the axes for a figure of its own: a legend inside sits on top of the data, and which corner is free depends on the run rather than on the plot - the figure would move its own key around as the numbers change. Beside it, the key is always in the same place and covers nothing.

A panel of someone else's figure is the exception. The room beside it belongs to the panel next to it, so a key anchored there is drawn over a neighbour rather than over the data; inside the panel it stays within the space the plot was given.

Parameters:

  • axes (Axes) –

    The axes the plot was drawn on.

  • handles (list[Any] | None, default: None ) –

    Legend handles, by default the ones already on the axes.

  • labels (list[str] | None, default: None ) –

    Their labels, by default the ones already on the axes.

note_missing(handles: list[Any], labels: list[str], missing: dict[tuple[str, str], int]) -> None

Add the legend entry that says how many values the figure could not draw.

What a plot can say beyond that depends on what it draws. A bar has a slot of its own to put a cross under, so BarPlot marks the categories themselves; a point in a cloud or a step of a sweep has no slot, and the count in the key is the whole statement there - enough that a filled value is not read as a measured one.

Parameters:

  • handles (list[Any]) –

    Legend handles, extended in place.

  • labels (list[str]) –

    Their labels, extended in place alongside handles.

  • missing (dict[tuple[str, str], int]) –

    Number of missing values per category and group, as counted by :meth:resolve_missing.

apply_theme() -> None

Install the seaborn theme this plot is drawn under, unless it has none.

The theme is matplotlib's global state rather than a property of one figure, so it is installed before the figure is built and left in place afterwards: a benchmark themes its plots by handing every one of them the same Theme, not by each plot putting the previous look back.

apply_grid(axes: Axes) -> None

Draw the gridlines the theme asks for, behind everything else on axes.

Parameters:

  • axes (Axes) –

    The axes the plot was drawn on.

setup_figure() -> None

Create a matplotlib figure, unless the plot is drawing into a shared axes.

save_figure(save_dir: str) -> list[Path]

Write the current figure to save_dir once per configured file format.

Parameters:

  • save_dir (str) –

    Directory to save the figure into. Created if it does not exist.

Returns:

  • list[Path] –

    Paths that were written successfully.

finalize_plot(xlabel: str, ylabel: str, title: str, ylim: tuple[float, float] | None = None, x_rotation: int = 45, save_dir: str | None = None) -> None

Apply common axis labels, title, limits, and display behavior.

Parameters:

  • xlabel (str) –

    Label for the x-axis.

  • ylabel (str) –

    Label for the y-axis.

  • title (str) –

    Plot title.

  • ylim (tuple[float, float] | None, default: None ) –

    Lower and upper y-axis limits, by default None.

  • x_rotation (int, default: 45 ) –

    Rotation angle for x-axis tick labels, by default 45.

  • save_dir (str | None, default: None ) –

    Directory to save the figure into, by default None.

apply_grouping(benchmark_results: BenchmarkResultContainer, rows: list[dict[str, Any]]) -> dict[str, Any]

Split rows into groups along :attr:grouping.

What that means is the grouper's business - a column of the plotted data, a value looked up per model, or a setting the algorithms were configured with - and so is deciding that it does not apply, in which case the bars stay ungrouped.

Parameters:

  • benchmark_results (BenchmarkResultContainer) –

    Benchmark data the feature results and algorithm configurations are read from.

  • rows (list[dict[str, Any]]) –

    Row-oriented plot data, annotated - and, where a grouping applies to only part of the data, reduced - in place.

Returns:

  • dict[str, Any] –

    Keyword arguments to forward to :meth:create. Empty when no grouping applies, so call sites can splat it unconditionally.

draw(*, benchmark_results: BenchmarkResultContainer, rows: list[dict[str, Any]], save_dir: str | None = None, **overrides: Any) -> None

Group rows and draw them with the display configuration of this plot.

This is what turns the declared fields - :attr:x, :attr:title, :attr:hline and the rest - into a :meth:create call, so a subclass only has to say which rows it plots. Doing it in one place is also what keeps :attr:group_by working for every bar plot rather than for those that remember to apply it.

Parameters:

  • benchmark_results (BenchmarkResultContainer) –

    Benchmark data, used to look up the groups of a feature :attr:group_by.

  • rows (list[dict[str, Any]]) –

    Row-oriented plot data.

  • save_dir (str | None, default: None ) –

    Directory to save the figure into, by default None.

  • **overrides (Any, default: {} ) –

    Keyword arguments forwarded to :meth:create, overriding the fields.

transform_rows(rows: list[dict[str, Any]], x: str | None, group: str | None) -> list[dict[str, Any]]

Return the rows to plot, by default the rows as they are.

A subclass that has to reduce its rows before they are drawn - pooling counts into a single ratio, say - overrides this rather than :meth:run, so it keeps the shared grouping and display handling. It is told what the bars and the groups turned out to be, since that is what a row has to keep to stay one of them.

Parameters:

  • rows (list[dict[str, Any]]) –

    Row-oriented plot data, already annotated with the dimensions' columns.

  • x (str | None) –

    Column the bars are drawn per, or None when the plot has no rows.

  • group (str | None) –

    Column the bars are split by, or None when they are ungrouped.

Returns:

create(*, rows: list[dict[str, Any]], xlabel: str, ylabel: str, title: str, x: str = 'x', y: str = 'y', aggregation: Aggregation = Aggregation.MEAN, errorbar: ErrorBar | str = AUTO_ERRORBAR, hue: str | None = None, hline: float | None = None, hline_label: str | None = None, hcolor: str = REFERENCE_LINE_COLOUR, baseline: float | None = None, ylim: tuple[float, float] | None = None, legend: bool = False, save_dir: str | None = None, **kwargs: Any) -> None

Create a bar plot from row-oriented data.

Parameters:

  • rows (dict[str, Any]) –

    Row-oriented mapping used to construct the plotting DataFrame.

  • xlabel (str) –

    Label for the x-axis.

  • ylabel (str) –

    Label for the y-axis.

  • title (str) –

    Plot title.

  • x (str, default: 'x' ) –

    Column name mapped to the x-axis, by default "x".

  • y (str, default: 'y' ) –

    Column name mapped to the y-axis, by default "y".

  • aggregation (Aggregation, default: MEAN ) –

    Aggregation strategy applied by seaborn, by default Aggregation.MEAN.

  • errorbar (ErrorBar | str, default: AUTO_ERRORBAR ) –

    Seaborn error bar specification ("sd", ("ci", 95), None to disable). By default "auto", which takes the error bar from aggregation: the spread of the samples for means, none for min/max.

  • hue (str | None, default: None ) –

    Optional grouping column for grouped bars, by default None.

  • hline (float | None, default: None ) –

    Optional horizontal reference line value, by default None.

  • hline_label (str | None, default: None ) –

    Legend label for the horizontal reference line, by default None.

  • hcolor (str, default: REFERENCE_LINE_COLOUR ) –

    Colour of the horizontal reference line, by default black.

  • baseline (float | None, default: None ) –

    Height of a solid black baseline marking where the bars start, by default None. Unlike hline it carries no label and stays out of the legend - it says where zero is, it does not name a target.

  • ylim (tuple[float, float] | None, default: None ) –

    Lower and upper y-axis limits, by default None.

  • legend (bool, default: False ) –

    Whether seaborn should create a legend for hue groups, by default False.

  • save_dir (str | None, default: None ) –

    Directory to save the figure into, by default None.

  • **kwargs (Any, default: {} ) –

    Additional keyword arguments forwarded to :func:seaborn.barplot. They override the defaults computed here, so anything seaborn understands (palette, saturation, capsize, err_kws, ...) can be tuned from the call site.

annotation_text(value: float) -> str

Return the text written above a bar of value.

Applies :attr:annotate_format, unless :attr:annotate_max_decimals allows the value to be written as a plain decimal instead of in scientific notation. A value read off a percent axis is written as a percentage, so the annotation says the same thing as the axis it stands on - unless a format was asked for, which wins.

Parameters:

  • value (float) –

    The aggregated value of one bar.

Returns:

  • str –

    The annotation, e.g. "0.000057" rather than "5.67e-05".

value(metric_result: MetricResult) -> float

Return the number a single metric result contributes.

Parameters:

  • metric_result (MetricResult) –

    One result of :attr:metric_cls.

Returns:

  • float –

    The value plotted for this result, by default the attribute :attr:y names.

rows(benchmark_results: BenchmarkResultContainer) -> list[dict[str, Any]]

Return one row per model and algorithm.

Parameters:

Returns:

  • list[dict[str, Any]] –

    Rows carrying the algorithm, the model, and the plotted value under :attr:y.

RuntimePlot

Bases: MetricBarPlot

Bar chart of the wall-clock runtime each algorithm needed.

One bar per algorithm, aggregated over every model in the benchmark - the mean by default, see :attr:aggregation - with the spread across those models as the error bar. Lower is better.

Requires the Runtime metric.

Every display option is inherited and can be set when the plot is constructed, e.g. RuntimePlot(annotate=False, file_formats=("pgf", "png")). BarPlot documents the colours, error bars, value annotations and grouping by a feature; SeabornPlot the figure size and the output formats.

Attributes:

  • figure_filename (str) –

    Stem of the written figure files, by default "runtime".

Examples:

>>> bench.add_metric(name="runtime", metric=Runtime())
>>> bench.add_plot(name="avg_runtime", plot=RuntimePlot())
See Also

RuntimePerModelPlot : The same numbers broken down per model.

y: MetricDimension = MetricDimension('runtime_seconds', 'Runtime (s)', baseline=0.0) class-attribute instance-attribute

Runtime in seconds, standing on a solid line at zero.

A runtime cannot go below zero, and a wide error bar reaching under the bars leaves them floating in the axes without something to stand on. The line is where the bars start, which is what tells a solver that took twice as long from one that took half.

theme: Theme | None = Theme() class-attribute instance-attribute

The seaborn theme the figure is drawn under, and the gridlines behind the marks.

A figure is read against its axis, so it is drawn with the lines that make that possible unless asked otherwise. None takes the theme and the grid away.

missing: Missing = Missing() class-attribute instance-attribute

What becomes of the values the plot cannot draw, and how it says they were there.

option_bundles: dict[str, type[OptionBundle]] = {'style': PlotStyle} class-attribute

Constructor arguments that configure several fields at once, and the bundle each takes.

A style is spread over the bundle fields rather than stored, so a benchmark can hand the same look to every plot while each keeps what it says itself. Applied least specific first: the shared style, then a bundle passed to the plot, then a flat option.

x: Dimension = AlgorithmDimension() class-attribute instance-attribute

What the bars are: one per value of this dimension, and its title on the axis.

aggregation: Aggregation = Aggregation.MEAN class-attribute instance-attribute

Aggregation applied to the values of an x category, by default their mean.

errorbars: ErrorBars | None = ErrorBars() class-attribute instance-attribute

The error bars drawn on top of the bars: what they show, their colour and caps.

None draws none, the same as ErrorBars(spec=None).

annotation: Annotation | None = None class-attribute instance-attribute

The values written above the bars - how they are formatted and how large.

None, the default, writes none: a bar chart is read off its axis, and a number above every bar is worth its clutter only when the exact value is the point. Pass an Annotation to turn them on, empty for the defaults.

grouping: Dimension | None = None class-attribute instance-attribute

What splits each bar into a group of bars.

One of the groupers - ModelDimension, AlgorithmDimension, FeatureDimension, ParameterDimension - or None, which leaves the bars ungrouped.

metric_cls: MetricClass property

The metric this plot reads, taken from its @plot(...) declaration.

Returns:

  • MetricClass –

    The first declared metric class.

Raises:

  • PlotMetricUndeclaredError –

    If the plot declares no metric, so there is nothing to read.

run(benchmark_results: BenchmarkResultContainer, save_dir: str | None = None) -> None

Generate plot output from benchmark results.

Parameters:

  • benchmark_results (BenchmarkResultContainer) –

    Aggregated benchmark data consumed by the plot implementation.

  • save_dir (str | None, default: None ) –

    Directory to save the figure into, by default None.

draw_into(axes: Axes, benchmark_results: BenchmarkResultContainer) -> None

Run this plot onto an existing axes rather than into a figure of its own.

Everything the plot draws goes through the pyplot state, so making that axes the current one is enough to redirect it. The figure, the files, and the window stay the caller's business - which is what lets several plots share one figure, e.g. the summary grid.

Parameters:

  • axes (Axes) –

    The axes to draw on.

  • benchmark_results (BenchmarkResultContainer) –

    Aggregated benchmark data handed to :meth:run.

resolve_missing(df: pd.DataFrame, column: str, *, by: str | None = None, within: str | None = None) -> tuple[pd.DataFrame, dict[tuple[str, str], int]]

Return the drawable rows and how many of them were not, per category.

A metric with nothing to report says so with a None or an infinity - a time to solution of a run that never reached the optimum is the usual one, since the expected time to something that did not happen is unbounded. Neither is a height a bar can have, and leaving them in poisons the aggregate: one infinity turns the mean of an algorithm into an infinity, and a missing value silently shortens it.

What happens to them is :attr:missing, and by default it is nothing: the plot raises rather than quietly showing a mean over fewer models than it claims. Asked to carry on, it either leaves them out or fills them from the values that could be drawn - Missing(policy="max") puts them just past the tallest bar, which is where "worse than everything here" belongs. Either way they are counted and a warning is logged: a bar resting on half its models is a different statement from one resting on all of them, and that is not visible in the bar itself.

Parameters:

  • df (DataFrame) –

    The plotting data.

  • column (str) –

    Column holding the plotted value.

  • by (str | None, default: None ) –

    Column whose categories the missing values are counted per, e.g. the x-axis of a bar plot. Without one they are counted under "".

  • within (str | None, default: None ) –

    Column that splits those categories further, e.g. the grouping of a bar plot. Counting per group is what lets a figure mark the one bar of a group that lost values rather than the whole category it sits in.

Returns:

  • tuple[DataFrame, dict[tuple[str, str], int]] –

    The rows to draw, and the number of missing values per category and group - the group is "" where there is none. The mapping is empty when nothing was missing.

Raises:

  • PlotMissingValuesError –

    If values are missing and the policy is "raise".

place_legend(axes: Axes, handles: list[Any] | None = None, labels: list[str] | None = None) -> None

Put the legend beside the axes, whatever drew it.

Outside the axes for a figure of its own: a legend inside sits on top of the data, and which corner is free depends on the run rather than on the plot - the figure would move its own key around as the numbers change. Beside it, the key is always in the same place and covers nothing.

A panel of someone else's figure is the exception. The room beside it belongs to the panel next to it, so a key anchored there is drawn over a neighbour rather than over the data; inside the panel it stays within the space the plot was given.

Parameters:

  • axes (Axes) –

    The axes the plot was drawn on.

  • handles (list[Any] | None, default: None ) –

    Legend handles, by default the ones already on the axes.

  • labels (list[str] | None, default: None ) –

    Their labels, by default the ones already on the axes.

note_missing(handles: list[Any], labels: list[str], missing: dict[tuple[str, str], int]) -> None

Add the legend entry that says how many values the figure could not draw.

What a plot can say beyond that depends on what it draws. A bar has a slot of its own to put a cross under, so BarPlot marks the categories themselves; a point in a cloud or a step of a sweep has no slot, and the count in the key is the whole statement there - enough that a filled value is not read as a measured one.

Parameters:

  • handles (list[Any]) –

    Legend handles, extended in place.

  • labels (list[str]) –

    Their labels, extended in place alongside handles.

  • missing (dict[tuple[str, str], int]) –

    Number of missing values per category and group, as counted by :meth:resolve_missing.

apply_theme() -> None

Install the seaborn theme this plot is drawn under, unless it has none.

The theme is matplotlib's global state rather than a property of one figure, so it is installed before the figure is built and left in place afterwards: a benchmark themes its plots by handing every one of them the same Theme, not by each plot putting the previous look back.

apply_grid(axes: Axes) -> None

Draw the gridlines the theme asks for, behind everything else on axes.

Parameters:

  • axes (Axes) –

    The axes the plot was drawn on.

setup_figure() -> None

Create a matplotlib figure, unless the plot is drawing into a shared axes.

save_figure(save_dir: str) -> list[Path]

Write the current figure to save_dir once per configured file format.

Parameters:

  • save_dir (str) –

    Directory to save the figure into. Created if it does not exist.

Returns:

  • list[Path] –

    Paths that were written successfully.

finalize_plot(xlabel: str, ylabel: str, title: str, ylim: tuple[float, float] | None = None, x_rotation: int = 45, save_dir: str | None = None) -> None

Apply common axis labels, title, limits, and display behavior.

Parameters:

  • xlabel (str) –

    Label for the x-axis.

  • ylabel (str) –

    Label for the y-axis.

  • title (str) –

    Plot title.

  • ylim (tuple[float, float] | None, default: None ) –

    Lower and upper y-axis limits, by default None.

  • x_rotation (int, default: 45 ) –

    Rotation angle for x-axis tick labels, by default 45.

  • save_dir (str | None, default: None ) –

    Directory to save the figure into, by default None.

apply_grouping(benchmark_results: BenchmarkResultContainer, rows: list[dict[str, Any]]) -> dict[str, Any]

Split rows into groups along :attr:grouping.

What that means is the grouper's business - a column of the plotted data, a value looked up per model, or a setting the algorithms were configured with - and so is deciding that it does not apply, in which case the bars stay ungrouped.

Parameters:

  • benchmark_results (BenchmarkResultContainer) –

    Benchmark data the feature results and algorithm configurations are read from.

  • rows (list[dict[str, Any]]) –

    Row-oriented plot data, annotated - and, where a grouping applies to only part of the data, reduced - in place.

Returns:

  • dict[str, Any] –

    Keyword arguments to forward to :meth:create. Empty when no grouping applies, so call sites can splat it unconditionally.

draw(*, benchmark_results: BenchmarkResultContainer, rows: list[dict[str, Any]], save_dir: str | None = None, **overrides: Any) -> None

Group rows and draw them with the display configuration of this plot.

This is what turns the declared fields - :attr:x, :attr:title, :attr:hline and the rest - into a :meth:create call, so a subclass only has to say which rows it plots. Doing it in one place is also what keeps :attr:group_by working for every bar plot rather than for those that remember to apply it.

Parameters:

  • benchmark_results (BenchmarkResultContainer) –

    Benchmark data, used to look up the groups of a feature :attr:group_by.

  • rows (list[dict[str, Any]]) –

    Row-oriented plot data.

  • save_dir (str | None, default: None ) –

    Directory to save the figure into, by default None.

  • **overrides (Any, default: {} ) –

    Keyword arguments forwarded to :meth:create, overriding the fields.

transform_rows(rows: list[dict[str, Any]], x: str | None, group: str | None) -> list[dict[str, Any]]

Return the rows to plot, by default the rows as they are.

A subclass that has to reduce its rows before they are drawn - pooling counts into a single ratio, say - overrides this rather than :meth:run, so it keeps the shared grouping and display handling. It is told what the bars and the groups turned out to be, since that is what a row has to keep to stay one of them.

Parameters:

  • rows (list[dict[str, Any]]) –

    Row-oriented plot data, already annotated with the dimensions' columns.

  • x (str | None) –

    Column the bars are drawn per, or None when the plot has no rows.

  • group (str | None) –

    Column the bars are split by, or None when they are ungrouped.

Returns:

create(*, rows: list[dict[str, Any]], xlabel: str, ylabel: str, title: str, x: str = 'x', y: str = 'y', aggregation: Aggregation = Aggregation.MEAN, errorbar: ErrorBar | str = AUTO_ERRORBAR, hue: str | None = None, hline: float | None = None, hline_label: str | None = None, hcolor: str = REFERENCE_LINE_COLOUR, baseline: float | None = None, ylim: tuple[float, float] | None = None, legend: bool = False, save_dir: str | None = None, **kwargs: Any) -> None

Create a bar plot from row-oriented data.

Parameters:

  • rows (dict[str, Any]) –

    Row-oriented mapping used to construct the plotting DataFrame.

  • xlabel (str) –

    Label for the x-axis.

  • ylabel (str) –

    Label for the y-axis.

  • title (str) –

    Plot title.

  • x (str, default: 'x' ) –

    Column name mapped to the x-axis, by default "x".

  • y (str, default: 'y' ) –

    Column name mapped to the y-axis, by default "y".

  • aggregation (Aggregation, default: MEAN ) –

    Aggregation strategy applied by seaborn, by default Aggregation.MEAN.

  • errorbar (ErrorBar | str, default: AUTO_ERRORBAR ) –

    Seaborn error bar specification ("sd", ("ci", 95), None to disable). By default "auto", which takes the error bar from aggregation: the spread of the samples for means, none for min/max.

  • hue (str | None, default: None ) –

    Optional grouping column for grouped bars, by default None.

  • hline (float | None, default: None ) –

    Optional horizontal reference line value, by default None.

  • hline_label (str | None, default: None ) –

    Legend label for the horizontal reference line, by default None.

  • hcolor (str, default: REFERENCE_LINE_COLOUR ) –

    Colour of the horizontal reference line, by default black.

  • baseline (float | None, default: None ) –

    Height of a solid black baseline marking where the bars start, by default None. Unlike hline it carries no label and stays out of the legend - it says where zero is, it does not name a target.

  • ylim (tuple[float, float] | None, default: None ) –

    Lower and upper y-axis limits, by default None.

  • legend (bool, default: False ) –

    Whether seaborn should create a legend for hue groups, by default False.

  • save_dir (str | None, default: None ) –

    Directory to save the figure into, by default None.

  • **kwargs (Any, default: {} ) –

    Additional keyword arguments forwarded to :func:seaborn.barplot. They override the defaults computed here, so anything seaborn understands (palette, saturation, capsize, err_kws, ...) can be tuned from the call site.

annotation_text(value: float) -> str

Return the text written above a bar of value.

Applies :attr:annotate_format, unless :attr:annotate_max_decimals allows the value to be written as a plain decimal instead of in scientific notation. A value read off a percent axis is written as a percentage, so the annotation says the same thing as the axis it stands on - unless a format was asked for, which wins.

Parameters:

  • value (float) –

    The aggregated value of one bar.

Returns:

  • str –

    The annotation, e.g. "0.000057" rather than "5.67e-05".

value(metric_result: MetricResult) -> float

Return the number a single metric result contributes.

Parameters:

  • metric_result (MetricResult) –

    One result of :attr:metric_cls.

Returns:

  • float –

    The value plotted for this result, by default the attribute :attr:y names.

rows(benchmark_results: BenchmarkResultContainer) -> list[dict[str, Any]]

Return one row per model and algorithm.

Parameters:

Returns:

  • list[dict[str, Any]] –

    Rows carrying the algorithm, the model, and the plotted value under :attr:y.

TimeToSolutionPlot

Bases: MetricBarPlot

Bar chart of the time to solution (TTS) per algorithm.

TTS is the runtime an algorithm needs to reach its target solution with a given confidence, so it weighs speed against success rate. One bar per algorithm, averaged over every model, with the spread across those models as the error bar. Lower is better.

Requires the TimeToSolution metric.

Every display option is inherited and can be set when the plot is constructed, e.g. TimeToSolutionPlot(annotate=False, file_formats=("pgf", "png")). BarPlot documents the colours, error bars, value annotations and grouping by a feature; SeabornPlot the figure size and the output formats.

Attributes:

  • figure_filename (str) –

    Stem of the written figure files, by default "time_to_solution".

Examples:

>>> bench.add_metric(name="time_to_solution", metric=TimeToSolution())
>>> bench.add_plot(name="avg_tts", plot=TimeToSolutionPlot())

y: MetricDimension = MetricDimension('time_to_solution', 'Time to Solution (TTS)', baseline=0.0) class-attribute instance-attribute

The expected time to the optimum, standing on a solid line at zero.

A time cannot go below zero, and the spread across models is wide enough here to send an error bar under the bars, which leaves them floating in the axes without a line to stand on. See RuntimePlot, which is read on the same scale.

theme: Theme | None = Theme() class-attribute instance-attribute

The seaborn theme the figure is drawn under, and the gridlines behind the marks.

A figure is read against its axis, so it is drawn with the lines that make that possible unless asked otherwise. None takes the theme and the grid away.

missing: Missing = Missing() class-attribute instance-attribute

What becomes of the values the plot cannot draw, and how it says they were there.

option_bundles: dict[str, type[OptionBundle]] = {'style': PlotStyle} class-attribute

Constructor arguments that configure several fields at once, and the bundle each takes.

A style is spread over the bundle fields rather than stored, so a benchmark can hand the same look to every plot while each keeps what it says itself. Applied least specific first: the shared style, then a bundle passed to the plot, then a flat option.

x: Dimension = AlgorithmDimension() class-attribute instance-attribute

What the bars are: one per value of this dimension, and its title on the axis.

aggregation: Aggregation = Aggregation.MEAN class-attribute instance-attribute

Aggregation applied to the values of an x category, by default their mean.

errorbars: ErrorBars | None = ErrorBars() class-attribute instance-attribute

The error bars drawn on top of the bars: what they show, their colour and caps.

None draws none, the same as ErrorBars(spec=None).

annotation: Annotation | None = None class-attribute instance-attribute

The values written above the bars - how they are formatted and how large.

None, the default, writes none: a bar chart is read off its axis, and a number above every bar is worth its clutter only when the exact value is the point. Pass an Annotation to turn them on, empty for the defaults.

grouping: Dimension | None = None class-attribute instance-attribute

What splits each bar into a group of bars.

One of the groupers - ModelDimension, AlgorithmDimension, FeatureDimension, ParameterDimension - or None, which leaves the bars ungrouped.

metric_cls: MetricClass property

The metric this plot reads, taken from its @plot(...) declaration.

Returns:

  • MetricClass –

    The first declared metric class.

Raises:

  • PlotMetricUndeclaredError –

    If the plot declares no metric, so there is nothing to read.

run(benchmark_results: BenchmarkResultContainer, save_dir: str | None = None) -> None

Generate plot output from benchmark results.

Parameters:

  • benchmark_results (BenchmarkResultContainer) –

    Aggregated benchmark data consumed by the plot implementation.

  • save_dir (str | None, default: None ) –

    Directory to save the figure into, by default None.

draw_into(axes: Axes, benchmark_results: BenchmarkResultContainer) -> None

Run this plot onto an existing axes rather than into a figure of its own.

Everything the plot draws goes through the pyplot state, so making that axes the current one is enough to redirect it. The figure, the files, and the window stay the caller's business - which is what lets several plots share one figure, e.g. the summary grid.

Parameters:

  • axes (Axes) –

    The axes to draw on.

  • benchmark_results (BenchmarkResultContainer) –

    Aggregated benchmark data handed to :meth:run.

resolve_missing(df: pd.DataFrame, column: str, *, by: str | None = None, within: str | None = None) -> tuple[pd.DataFrame, dict[tuple[str, str], int]]

Return the drawable rows and how many of them were not, per category.

A metric with nothing to report says so with a None or an infinity - a time to solution of a run that never reached the optimum is the usual one, since the expected time to something that did not happen is unbounded. Neither is a height a bar can have, and leaving them in poisons the aggregate: one infinity turns the mean of an algorithm into an infinity, and a missing value silently shortens it.

What happens to them is :attr:missing, and by default it is nothing: the plot raises rather than quietly showing a mean over fewer models than it claims. Asked to carry on, it either leaves them out or fills them from the values that could be drawn - Missing(policy="max") puts them just past the tallest bar, which is where "worse than everything here" belongs. Either way they are counted and a warning is logged: a bar resting on half its models is a different statement from one resting on all of them, and that is not visible in the bar itself.

Parameters:

  • df (DataFrame) –

    The plotting data.

  • column (str) –

    Column holding the plotted value.

  • by (str | None, default: None ) –

    Column whose categories the missing values are counted per, e.g. the x-axis of a bar plot. Without one they are counted under "".

  • within (str | None, default: None ) –

    Column that splits those categories further, e.g. the grouping of a bar plot. Counting per group is what lets a figure mark the one bar of a group that lost values rather than the whole category it sits in.

Returns:

  • tuple[DataFrame, dict[tuple[str, str], int]] –

    The rows to draw, and the number of missing values per category and group - the group is "" where there is none. The mapping is empty when nothing was missing.

Raises:

  • PlotMissingValuesError –

    If values are missing and the policy is "raise".

place_legend(axes: Axes, handles: list[Any] | None = None, labels: list[str] | None = None) -> None

Put the legend beside the axes, whatever drew it.

Outside the axes for a figure of its own: a legend inside sits on top of the data, and which corner is free depends on the run rather than on the plot - the figure would move its own key around as the numbers change. Beside it, the key is always in the same place and covers nothing.

A panel of someone else's figure is the exception. The room beside it belongs to the panel next to it, so a key anchored there is drawn over a neighbour rather than over the data; inside the panel it stays within the space the plot was given.

Parameters:

  • axes (Axes) –

    The axes the plot was drawn on.

  • handles (list[Any] | None, default: None ) –

    Legend handles, by default the ones already on the axes.

  • labels (list[str] | None, default: None ) –

    Their labels, by default the ones already on the axes.

note_missing(handles: list[Any], labels: list[str], missing: dict[tuple[str, str], int]) -> None

Add the legend entry that says how many values the figure could not draw.

What a plot can say beyond that depends on what it draws. A bar has a slot of its own to put a cross under, so BarPlot marks the categories themselves; a point in a cloud or a step of a sweep has no slot, and the count in the key is the whole statement there - enough that a filled value is not read as a measured one.

Parameters:

  • handles (list[Any]) –

    Legend handles, extended in place.

  • labels (list[str]) –

    Their labels, extended in place alongside handles.

  • missing (dict[tuple[str, str], int]) –

    Number of missing values per category and group, as counted by :meth:resolve_missing.

apply_theme() -> None

Install the seaborn theme this plot is drawn under, unless it has none.

The theme is matplotlib's global state rather than a property of one figure, so it is installed before the figure is built and left in place afterwards: a benchmark themes its plots by handing every one of them the same Theme, not by each plot putting the previous look back.

apply_grid(axes: Axes) -> None

Draw the gridlines the theme asks for, behind everything else on axes.

Parameters:

  • axes (Axes) –

    The axes the plot was drawn on.

setup_figure() -> None

Create a matplotlib figure, unless the plot is drawing into a shared axes.

save_figure(save_dir: str) -> list[Path]

Write the current figure to save_dir once per configured file format.

Parameters:

  • save_dir (str) –

    Directory to save the figure into. Created if it does not exist.

Returns:

  • list[Path] –

    Paths that were written successfully.

finalize_plot(xlabel: str, ylabel: str, title: str, ylim: tuple[float, float] | None = None, x_rotation: int = 45, save_dir: str | None = None) -> None

Apply common axis labels, title, limits, and display behavior.

Parameters:

  • xlabel (str) –

    Label for the x-axis.

  • ylabel (str) –

    Label for the y-axis.

  • title (str) –

    Plot title.

  • ylim (tuple[float, float] | None, default: None ) –

    Lower and upper y-axis limits, by default None.

  • x_rotation (int, default: 45 ) –

    Rotation angle for x-axis tick labels, by default 45.

  • save_dir (str | None, default: None ) –

    Directory to save the figure into, by default None.

apply_grouping(benchmark_results: BenchmarkResultContainer, rows: list[dict[str, Any]]) -> dict[str, Any]

Split rows into groups along :attr:grouping.

What that means is the grouper's business - a column of the plotted data, a value looked up per model, or a setting the algorithms were configured with - and so is deciding that it does not apply, in which case the bars stay ungrouped.

Parameters:

  • benchmark_results (BenchmarkResultContainer) –

    Benchmark data the feature results and algorithm configurations are read from.

  • rows (list[dict[str, Any]]) –

    Row-oriented plot data, annotated - and, where a grouping applies to only part of the data, reduced - in place.

Returns:

  • dict[str, Any] –

    Keyword arguments to forward to :meth:create. Empty when no grouping applies, so call sites can splat it unconditionally.

draw(*, benchmark_results: BenchmarkResultContainer, rows: list[dict[str, Any]], save_dir: str | None = None, **overrides: Any) -> None

Group rows and draw them with the display configuration of this plot.

This is what turns the declared fields - :attr:x, :attr:title, :attr:hline and the rest - into a :meth:create call, so a subclass only has to say which rows it plots. Doing it in one place is also what keeps :attr:group_by working for every bar plot rather than for those that remember to apply it.

Parameters:

  • benchmark_results (BenchmarkResultContainer) –

    Benchmark data, used to look up the groups of a feature :attr:group_by.

  • rows (list[dict[str, Any]]) –

    Row-oriented plot data.

  • save_dir (str | None, default: None ) –

    Directory to save the figure into, by default None.

  • **overrides (Any, default: {} ) –

    Keyword arguments forwarded to :meth:create, overriding the fields.

transform_rows(rows: list[dict[str, Any]], x: str | None, group: str | None) -> list[dict[str, Any]]

Return the rows to plot, by default the rows as they are.

A subclass that has to reduce its rows before they are drawn - pooling counts into a single ratio, say - overrides this rather than :meth:run, so it keeps the shared grouping and display handling. It is told what the bars and the groups turned out to be, since that is what a row has to keep to stay one of them.

Parameters:

  • rows (list[dict[str, Any]]) –

    Row-oriented plot data, already annotated with the dimensions' columns.

  • x (str | None) –

    Column the bars are drawn per, or None when the plot has no rows.

  • group (str | None) –

    Column the bars are split by, or None when they are ungrouped.

Returns:

create(*, rows: list[dict[str, Any]], xlabel: str, ylabel: str, title: str, x: str = 'x', y: str = 'y', aggregation: Aggregation = Aggregation.MEAN, errorbar: ErrorBar | str = AUTO_ERRORBAR, hue: str | None = None, hline: float | None = None, hline_label: str | None = None, hcolor: str = REFERENCE_LINE_COLOUR, baseline: float | None = None, ylim: tuple[float, float] | None = None, legend: bool = False, save_dir: str | None = None, **kwargs: Any) -> None

Create a bar plot from row-oriented data.

Parameters:

  • rows (dict[str, Any]) –

    Row-oriented mapping used to construct the plotting DataFrame.

  • xlabel (str) –

    Label for the x-axis.

  • ylabel (str) –

    Label for the y-axis.

  • title (str) –

    Plot title.

  • x (str, default: 'x' ) –

    Column name mapped to the x-axis, by default "x".

  • y (str, default: 'y' ) –

    Column name mapped to the y-axis, by default "y".

  • aggregation (Aggregation, default: MEAN ) –

    Aggregation strategy applied by seaborn, by default Aggregation.MEAN.

  • errorbar (ErrorBar | str, default: AUTO_ERRORBAR ) –

    Seaborn error bar specification ("sd", ("ci", 95), None to disable). By default "auto", which takes the error bar from aggregation: the spread of the samples for means, none for min/max.

  • hue (str | None, default: None ) –

    Optional grouping column for grouped bars, by default None.

  • hline (float | None, default: None ) –

    Optional horizontal reference line value, by default None.

  • hline_label (str | None, default: None ) –

    Legend label for the horizontal reference line, by default None.

  • hcolor (str, default: REFERENCE_LINE_COLOUR ) –

    Colour of the horizontal reference line, by default black.

  • baseline (float | None, default: None ) –

    Height of a solid black baseline marking where the bars start, by default None. Unlike hline it carries no label and stays out of the legend - it says where zero is, it does not name a target.

  • ylim (tuple[float, float] | None, default: None ) –

    Lower and upper y-axis limits, by default None.

  • legend (bool, default: False ) –

    Whether seaborn should create a legend for hue groups, by default False.

  • save_dir (str | None, default: None ) –

    Directory to save the figure into, by default None.

  • **kwargs (Any, default: {} ) –

    Additional keyword arguments forwarded to :func:seaborn.barplot. They override the defaults computed here, so anything seaborn understands (palette, saturation, capsize, err_kws, ...) can be tuned from the call site.

annotation_text(value: float) -> str

Return the text written above a bar of value.

Applies :attr:annotate_format, unless :attr:annotate_max_decimals allows the value to be written as a plain decimal instead of in scientific notation. A value read off a percent axis is written as a percentage, so the annotation says the same thing as the axis it stands on - unless a format was asked for, which wins.

Parameters:

  • value (float) –

    The aggregated value of one bar.

Returns:

  • str –

    The annotation, e.g. "0.000057" rather than "5.67e-05".

value(metric_result: MetricResult) -> float

Return the number a single metric result contributes.

Parameters:

  • metric_result (MetricResult) –

    One result of :attr:metric_cls.

Returns:

  • float –

    The value plotted for this result, by default the attribute :attr:y names.

rows(benchmark_results: BenchmarkResultContainer) -> list[dict[str, Any]]

Return one row per model and algorithm.

Parameters:

Returns:

  • list[dict[str, Any]] –

    Rows carrying the algorithm, the model, and the plotted value under :attr:y.

Analysis Plots

ApproximationRatioVsVarNumberPlot

Bases: ScatterPlot

Scatter plot showing an approximation ratio vs. number of variables per model/algorithm.

Examples:

>>> bench.add_feature(name="var_count", feature=VarNumberFeature())
>>> bench.add_metric(name="approx_ratio", metric=ApproximationRatio())
>>> bench.add_plot(name="approx_vs_vars", plot=ApproximationRatioVsVarNumberPlot())

theme: Theme | None = Theme() class-attribute instance-attribute

The seaborn theme the figure is drawn under, and the gridlines behind the marks.

A figure is read against its axis, so it is drawn with the lines that make that possible unless asked otherwise. None takes the theme and the grid away.

missing: Missing = Missing() class-attribute instance-attribute

What becomes of the values the plot cannot draw, and how it says they were there.

option_bundles: dict[str, type[OptionBundle]] = {'style': PlotStyle} class-attribute

Constructor arguments that configure several fields at once, and the bundle each takes.

A style is spread over the bundle fields rather than stored, so a benchmark can hand the same look to every plot while each keeps what it says itself. Applied least specific first: the shared style, then a bundle passed to the plot, then a flat option.

run(benchmark_results: BenchmarkResultContainer, save_dir: str | None = None) -> None

Generate plot output from benchmark results.

Parameters:

draw_into(axes: Axes, benchmark_results: BenchmarkResultContainer) -> None

Run this plot onto an existing axes rather than into a figure of its own.

Everything the plot draws goes through the pyplot state, so making that axes the current one is enough to redirect it. The figure, the files, and the window stay the caller's business - which is what lets several plots share one figure, e.g. the summary grid.

Parameters:

  • axes (Axes) –

    The axes to draw on.

  • benchmark_results (BenchmarkResultContainer) –

    Aggregated benchmark data handed to :meth:run.

resolve_missing(df: pd.DataFrame, column: str, *, by: str | None = None, within: str | None = None) -> tuple[pd.DataFrame, dict[tuple[str, str], int]]

Return the drawable rows and how many of them were not, per category.

A metric with nothing to report says so with a None or an infinity - a time to solution of a run that never reached the optimum is the usual one, since the expected time to something that did not happen is unbounded. Neither is a height a bar can have, and leaving them in poisons the aggregate: one infinity turns the mean of an algorithm into an infinity, and a missing value silently shortens it.

What happens to them is :attr:missing, and by default it is nothing: the plot raises rather than quietly showing a mean over fewer models than it claims. Asked to carry on, it either leaves them out or fills them from the values that could be drawn - Missing(policy="max") puts them just past the tallest bar, which is where "worse than everything here" belongs. Either way they are counted and a warning is logged: a bar resting on half its models is a different statement from one resting on all of them, and that is not visible in the bar itself.

Parameters:

  • df (DataFrame) –

    The plotting data.

  • column (str) –

    Column holding the plotted value.

  • by (str | None, default: None ) –

    Column whose categories the missing values are counted per, e.g. the x-axis of a bar plot. Without one they are counted under "".

  • within (str | None, default: None ) –

    Column that splits those categories further, e.g. the grouping of a bar plot. Counting per group is what lets a figure mark the one bar of a group that lost values rather than the whole category it sits in.

Returns:

  • tuple[DataFrame, dict[tuple[str, str], int]] –

    The rows to draw, and the number of missing values per category and group - the group is "" where there is none. The mapping is empty when nothing was missing.

Raises:

  • PlotMissingValuesError –

    If values are missing and the policy is "raise".

place_legend(axes: Axes, handles: list[Any] | None = None, labels: list[str] | None = None) -> None

Put the legend beside the axes, whatever drew it.

Outside the axes for a figure of its own: a legend inside sits on top of the data, and which corner is free depends on the run rather than on the plot - the figure would move its own key around as the numbers change. Beside it, the key is always in the same place and covers nothing.

A panel of someone else's figure is the exception. The room beside it belongs to the panel next to it, so a key anchored there is drawn over a neighbour rather than over the data; inside the panel it stays within the space the plot was given.

Parameters:

  • axes (Axes) –

    The axes the plot was drawn on.

  • handles (list[Any] | None, default: None ) –

    Legend handles, by default the ones already on the axes.

  • labels (list[str] | None, default: None ) –

    Their labels, by default the ones already on the axes.

note_missing(handles: list[Any], labels: list[str], missing: dict[tuple[str, str], int]) -> None

Add the legend entry that says how many values the figure could not draw.

What a plot can say beyond that depends on what it draws. A bar has a slot of its own to put a cross under, so BarPlot marks the categories themselves; a point in a cloud or a step of a sweep has no slot, and the count in the key is the whole statement there - enough that a filled value is not read as a measured one.

Parameters:

  • handles (list[Any]) –

    Legend handles, extended in place.

  • labels (list[str]) –

    Their labels, extended in place alongside handles.

  • missing (dict[tuple[str, str], int]) –

    Number of missing values per category and group, as counted by :meth:resolve_missing.

apply_theme() -> None

Install the seaborn theme this plot is drawn under, unless it has none.

The theme is matplotlib's global state rather than a property of one figure, so it is installed before the figure is built and left in place afterwards: a benchmark themes its plots by handing every one of them the same Theme, not by each plot putting the previous look back.

apply_grid(axes: Axes) -> None

Draw the gridlines the theme asks for, behind everything else on axes.

Parameters:

  • axes (Axes) –

    The axes the plot was drawn on.

setup_figure() -> None

Create a matplotlib figure, unless the plot is drawing into a shared axes.

save_figure(save_dir: str) -> list[Path]

Write the current figure to save_dir once per configured file format.

Parameters:

  • save_dir (str) –

    Directory to save the figure into. Created if it does not exist.

Returns:

  • list[Path] –

    Paths that were written successfully.

finalize_plot(xlabel: str, ylabel: str, title: str, ylim: tuple[float, float] | None = None, x_rotation: int = 45, save_dir: str | None = None) -> None

Apply common axis labels, title, limits, and display behavior.

Parameters:

  • xlabel (str) –

    Label for the x-axis.

  • ylabel (str) –

    Label for the y-axis.

  • title (str) –

    Plot title.

  • ylim (tuple[float, float] | None, default: None ) –

    Lower and upper y-axis limits, by default None.

  • x_rotation (int, default: 45 ) –

    Rotation angle for x-axis tick labels, by default 45.

  • save_dir (str | None, default: None ) –

    Directory to save the figure into, by default None.

create(*, rows: list[dict[str, Any]], xlabel: str, ylabel: str, title: str, hue: str, x: str = 'x', y: str = 'y', hline: float | None = None, hline_label: str | None = None, hcolor: str = REFERENCE_LINE_COLOUR, save_dir: str | None = None, **kwargs: Any) -> None

Create a scatter plot from row-oriented data.

Parameters:

  • rows (dict[str, Any]) –

    Row-oriented mapping used to construct the plotting DataFrame.

  • xlabel (str) –

    Label for the x-axis.

  • ylabel (str) –

    Label for the y-axis.

  • title (str) –

    Plot title.

  • hue (str) –

    Column used to color points by group.

  • x (str, default: 'x' ) –

    Column name mapped to the x-axis, by default "x".

  • y (str, default: 'y' ) –

    Column name mapped to the y-axis, by default "y".

  • hline (float | None, default: None ) –

    Optional horizontal reference line value, by default None.

  • hline_label (str | None, default: None ) –

    Legend label for the horizontal reference line, by default None.

  • hcolor (str, default: REFERENCE_LINE_COLOUR ) –

    Color of the horizontal reference line, by default black.

  • save_dir (str | None, default: None ) –

    Directory to save the figure into, by default None.

  • **kwargs (Any, default: {} ) –

    Additional keyword arguments forwarded to :func:seaborn.scatterplot. They override the defaults computed here, so anything seaborn understands (palette, style, size, markers, ...) can be tuned from the call site.

FeasibilityRatioVsVarNumberPlot

Bases: ScatterPlot

Scatter plot showing feasibility ratio vs number of variables per model/algorithm.

Examples:

>>> bench.add_feature(name="var_count", feature=VarNumberFeature())
>>> bench.add_metric(name="feasibility", metric=FeasibilityRatio())
>>> bench.add_plot(name="feasibility_vs_vars", plot=FeasibilityRatioVsVarNumberPlot())

theme: Theme | None = Theme() class-attribute instance-attribute

The seaborn theme the figure is drawn under, and the gridlines behind the marks.

A figure is read against its axis, so it is drawn with the lines that make that possible unless asked otherwise. None takes the theme and the grid away.

missing: Missing = Missing() class-attribute instance-attribute

What becomes of the values the plot cannot draw, and how it says they were there.

option_bundles: dict[str, type[OptionBundle]] = {'style': PlotStyle} class-attribute

Constructor arguments that configure several fields at once, and the bundle each takes.

A style is spread over the bundle fields rather than stored, so a benchmark can hand the same look to every plot while each keeps what it says itself. Applied least specific first: the shared style, then a bundle passed to the plot, then a flat option.

run(benchmark_results: BenchmarkResultContainer, save_dir: str | None = None) -> None

Generate plot output from benchmark results.

Parameters:

draw_into(axes: Axes, benchmark_results: BenchmarkResultContainer) -> None

Run this plot onto an existing axes rather than into a figure of its own.

Everything the plot draws goes through the pyplot state, so making that axes the current one is enough to redirect it. The figure, the files, and the window stay the caller's business - which is what lets several plots share one figure, e.g. the summary grid.

Parameters:

  • axes (Axes) –

    The axes to draw on.

  • benchmark_results (BenchmarkResultContainer) –

    Aggregated benchmark data handed to :meth:run.

resolve_missing(df: pd.DataFrame, column: str, *, by: str | None = None, within: str | None = None) -> tuple[pd.DataFrame, dict[tuple[str, str], int]]

Return the drawable rows and how many of them were not, per category.

A metric with nothing to report says so with a None or an infinity - a time to solution of a run that never reached the optimum is the usual one, since the expected time to something that did not happen is unbounded. Neither is a height a bar can have, and leaving them in poisons the aggregate: one infinity turns the mean of an algorithm into an infinity, and a missing value silently shortens it.

What happens to them is :attr:missing, and by default it is nothing: the plot raises rather than quietly showing a mean over fewer models than it claims. Asked to carry on, it either leaves them out or fills them from the values that could be drawn - Missing(policy="max") puts them just past the tallest bar, which is where "worse than everything here" belongs. Either way they are counted and a warning is logged: a bar resting on half its models is a different statement from one resting on all of them, and that is not visible in the bar itself.

Parameters:

  • df (DataFrame) –

    The plotting data.

  • column (str) –

    Column holding the plotted value.

  • by (str | None, default: None ) –

    Column whose categories the missing values are counted per, e.g. the x-axis of a bar plot. Without one they are counted under "".

  • within (str | None, default: None ) –

    Column that splits those categories further, e.g. the grouping of a bar plot. Counting per group is what lets a figure mark the one bar of a group that lost values rather than the whole category it sits in.

Returns:

  • tuple[DataFrame, dict[tuple[str, str], int]] –

    The rows to draw, and the number of missing values per category and group - the group is "" where there is none. The mapping is empty when nothing was missing.

Raises:

  • PlotMissingValuesError –

    If values are missing and the policy is "raise".

place_legend(axes: Axes, handles: list[Any] | None = None, labels: list[str] | None = None) -> None

Put the legend beside the axes, whatever drew it.

Outside the axes for a figure of its own: a legend inside sits on top of the data, and which corner is free depends on the run rather than on the plot - the figure would move its own key around as the numbers change. Beside it, the key is always in the same place and covers nothing.

A panel of someone else's figure is the exception. The room beside it belongs to the panel next to it, so a key anchored there is drawn over a neighbour rather than over the data; inside the panel it stays within the space the plot was given.

Parameters:

  • axes (Axes) –

    The axes the plot was drawn on.

  • handles (list[Any] | None, default: None ) –

    Legend handles, by default the ones already on the axes.

  • labels (list[str] | None, default: None ) –

    Their labels, by default the ones already on the axes.

note_missing(handles: list[Any], labels: list[str], missing: dict[tuple[str, str], int]) -> None

Add the legend entry that says how many values the figure could not draw.

What a plot can say beyond that depends on what it draws. A bar has a slot of its own to put a cross under, so BarPlot marks the categories themselves; a point in a cloud or a step of a sweep has no slot, and the count in the key is the whole statement there - enough that a filled value is not read as a measured one.

Parameters:

  • handles (list[Any]) –

    Legend handles, extended in place.

  • labels (list[str]) –

    Their labels, extended in place alongside handles.

  • missing (dict[tuple[str, str], int]) –

    Number of missing values per category and group, as counted by :meth:resolve_missing.

apply_theme() -> None

Install the seaborn theme this plot is drawn under, unless it has none.

The theme is matplotlib's global state rather than a property of one figure, so it is installed before the figure is built and left in place afterwards: a benchmark themes its plots by handing every one of them the same Theme, not by each plot putting the previous look back.

apply_grid(axes: Axes) -> None

Draw the gridlines the theme asks for, behind everything else on axes.

Parameters:

  • axes (Axes) –

    The axes the plot was drawn on.

setup_figure() -> None

Create a matplotlib figure, unless the plot is drawing into a shared axes.

save_figure(save_dir: str) -> list[Path]

Write the current figure to save_dir once per configured file format.

Parameters:

  • save_dir (str) –

    Directory to save the figure into. Created if it does not exist.

Returns:

  • list[Path] –

    Paths that were written successfully.

finalize_plot(xlabel: str, ylabel: str, title: str, ylim: tuple[float, float] | None = None, x_rotation: int = 45, save_dir: str | None = None) -> None

Apply common axis labels, title, limits, and display behavior.

Parameters:

  • xlabel (str) –

    Label for the x-axis.

  • ylabel (str) –

    Label for the y-axis.

  • title (str) –

    Plot title.

  • ylim (tuple[float, float] | None, default: None ) –

    Lower and upper y-axis limits, by default None.

  • x_rotation (int, default: 45 ) –

    Rotation angle for x-axis tick labels, by default 45.

  • save_dir (str | None, default: None ) –

    Directory to save the figure into, by default None.

create(*, rows: list[dict[str, Any]], xlabel: str, ylabel: str, title: str, hue: str, x: str = 'x', y: str = 'y', hline: float | None = None, hline_label: str | None = None, hcolor: str = REFERENCE_LINE_COLOUR, save_dir: str | None = None, **kwargs: Any) -> None

Create a scatter plot from row-oriented data.

Parameters:

  • rows (dict[str, Any]) –

    Row-oriented mapping used to construct the plotting DataFrame.

  • xlabel (str) –

    Label for the x-axis.

  • ylabel (str) –

    Label for the y-axis.

  • title (str) –

    Plot title.

  • hue (str) –

    Column used to color points by group.

  • x (str, default: 'x' ) –

    Column name mapped to the x-axis, by default "x".

  • y (str, default: 'y' ) –

    Column name mapped to the y-axis, by default "y".

  • hline (float | None, default: None ) –

    Optional horizontal reference line value, by default None.

  • hline_label (str | None, default: None ) –

    Legend label for the horizontal reference line, by default None.

  • hcolor (str, default: REFERENCE_LINE_COLOUR ) –

    Color of the horizontal reference line, by default black.

  • save_dir (str | None, default: None ) –

    Directory to save the figure into, by default None.

  • **kwargs (Any, default: {} ) –

    Additional keyword arguments forwarded to :func:seaborn.scatterplot. They override the defaults computed here, so anything seaborn understands (palette, style, size, markers, ...) can be tuned from the call site.

RuntimeVsVarNumberPlot

Bases: ScatterPlot

Scatter plot showing runtime vs number of variables per model/algorithm.

Examples:

>>> bench.add_feature(name="var_count", feature=VarNumberFeature())
>>> bench.add_metric(name="runtime", metric=Runtime())
>>> bench.add_plot(name="runtime_vs_vars", plot=RuntimeVsVarNumberPlot())

theme: Theme | None = Theme() class-attribute instance-attribute

The seaborn theme the figure is drawn under, and the gridlines behind the marks.

A figure is read against its axis, so it is drawn with the lines that make that possible unless asked otherwise. None takes the theme and the grid away.

missing: Missing = Missing() class-attribute instance-attribute

What becomes of the values the plot cannot draw, and how it says they were there.

option_bundles: dict[str, type[OptionBundle]] = {'style': PlotStyle} class-attribute

Constructor arguments that configure several fields at once, and the bundle each takes.

A style is spread over the bundle fields rather than stored, so a benchmark can hand the same look to every plot while each keeps what it says itself. Applied least specific first: the shared style, then a bundle passed to the plot, then a flat option.

run(benchmark_results: BenchmarkResultContainer, save_dir: str | None = None) -> None

Generate plot output from benchmark results.

Parameters:

draw_into(axes: Axes, benchmark_results: BenchmarkResultContainer) -> None

Run this plot onto an existing axes rather than into a figure of its own.

Everything the plot draws goes through the pyplot state, so making that axes the current one is enough to redirect it. The figure, the files, and the window stay the caller's business - which is what lets several plots share one figure, e.g. the summary grid.

Parameters:

  • axes (Axes) –

    The axes to draw on.

  • benchmark_results (BenchmarkResultContainer) –

    Aggregated benchmark data handed to :meth:run.

resolve_missing(df: pd.DataFrame, column: str, *, by: str | None = None, within: str | None = None) -> tuple[pd.DataFrame, dict[tuple[str, str], int]]

Return the drawable rows and how many of them were not, per category.

A metric with nothing to report says so with a None or an infinity - a time to solution of a run that never reached the optimum is the usual one, since the expected time to something that did not happen is unbounded. Neither is a height a bar can have, and leaving them in poisons the aggregate: one infinity turns the mean of an algorithm into an infinity, and a missing value silently shortens it.

What happens to them is :attr:missing, and by default it is nothing: the plot raises rather than quietly showing a mean over fewer models than it claims. Asked to carry on, it either leaves them out or fills them from the values that could be drawn - Missing(policy="max") puts them just past the tallest bar, which is where "worse than everything here" belongs. Either way they are counted and a warning is logged: a bar resting on half its models is a different statement from one resting on all of them, and that is not visible in the bar itself.

Parameters:

  • df (DataFrame) –

    The plotting data.

  • column (str) –

    Column holding the plotted value.

  • by (str | None, default: None ) –

    Column whose categories the missing values are counted per, e.g. the x-axis of a bar plot. Without one they are counted under "".

  • within (str | None, default: None ) –

    Column that splits those categories further, e.g. the grouping of a bar plot. Counting per group is what lets a figure mark the one bar of a group that lost values rather than the whole category it sits in.

Returns:

  • tuple[DataFrame, dict[tuple[str, str], int]] –

    The rows to draw, and the number of missing values per category and group - the group is "" where there is none. The mapping is empty when nothing was missing.

Raises:

  • PlotMissingValuesError –

    If values are missing and the policy is "raise".

place_legend(axes: Axes, handles: list[Any] | None = None, labels: list[str] | None = None) -> None

Put the legend beside the axes, whatever drew it.

Outside the axes for a figure of its own: a legend inside sits on top of the data, and which corner is free depends on the run rather than on the plot - the figure would move its own key around as the numbers change. Beside it, the key is always in the same place and covers nothing.

A panel of someone else's figure is the exception. The room beside it belongs to the panel next to it, so a key anchored there is drawn over a neighbour rather than over the data; inside the panel it stays within the space the plot was given.

Parameters:

  • axes (Axes) –

    The axes the plot was drawn on.

  • handles (list[Any] | None, default: None ) –

    Legend handles, by default the ones already on the axes.

  • labels (list[str] | None, default: None ) –

    Their labels, by default the ones already on the axes.

note_missing(handles: list[Any], labels: list[str], missing: dict[tuple[str, str], int]) -> None

Add the legend entry that says how many values the figure could not draw.

What a plot can say beyond that depends on what it draws. A bar has a slot of its own to put a cross under, so BarPlot marks the categories themselves; a point in a cloud or a step of a sweep has no slot, and the count in the key is the whole statement there - enough that a filled value is not read as a measured one.

Parameters:

  • handles (list[Any]) –

    Legend handles, extended in place.

  • labels (list[str]) –

    Their labels, extended in place alongside handles.

  • missing (dict[tuple[str, str], int]) –

    Number of missing values per category and group, as counted by :meth:resolve_missing.

apply_theme() -> None

Install the seaborn theme this plot is drawn under, unless it has none.

The theme is matplotlib's global state rather than a property of one figure, so it is installed before the figure is built and left in place afterwards: a benchmark themes its plots by handing every one of them the same Theme, not by each plot putting the previous look back.

apply_grid(axes: Axes) -> None

Draw the gridlines the theme asks for, behind everything else on axes.

Parameters:

  • axes (Axes) –

    The axes the plot was drawn on.

setup_figure() -> None

Create a matplotlib figure, unless the plot is drawing into a shared axes.

save_figure(save_dir: str) -> list[Path]

Write the current figure to save_dir once per configured file format.

Parameters:

  • save_dir (str) –

    Directory to save the figure into. Created if it does not exist.

Returns:

  • list[Path] –

    Paths that were written successfully.

finalize_plot(xlabel: str, ylabel: str, title: str, ylim: tuple[float, float] | None = None, x_rotation: int = 45, save_dir: str | None = None) -> None

Apply common axis labels, title, limits, and display behavior.

Parameters:

  • xlabel (str) –

    Label for the x-axis.

  • ylabel (str) –

    Label for the y-axis.

  • title (str) –

    Plot title.

  • ylim (tuple[float, float] | None, default: None ) –

    Lower and upper y-axis limits, by default None.

  • x_rotation (int, default: 45 ) –

    Rotation angle for x-axis tick labels, by default 45.

  • save_dir (str | None, default: None ) –

    Directory to save the figure into, by default None.

create(*, rows: list[dict[str, Any]], xlabel: str, ylabel: str, title: str, hue: str, x: str = 'x', y: str = 'y', hline: float | None = None, hline_label: str | None = None, hcolor: str = REFERENCE_LINE_COLOUR, save_dir: str | None = None, **kwargs: Any) -> None

Create a scatter plot from row-oriented data.

Parameters:

  • rows (dict[str, Any]) –

    Row-oriented mapping used to construct the plotting DataFrame.

  • xlabel (str) –

    Label for the x-axis.

  • ylabel (str) –

    Label for the y-axis.

  • title (str) –

    Plot title.

  • hue (str) –

    Column used to color points by group.

  • x (str, default: 'x' ) –

    Column name mapped to the x-axis, by default "x".

  • y (str, default: 'y' ) –

    Column name mapped to the y-axis, by default "y".

  • hline (float | None, default: None ) –

    Optional horizontal reference line value, by default None.

  • hline_label (str | None, default: None ) –

    Legend label for the horizontal reference line, by default None.

  • hcolor (str, default: REFERENCE_LINE_COLOUR ) –

    Color of the horizontal reference line, by default black.

  • save_dir (str | None, default: None ) –

    Directory to save the figure into, by default None.

  • **kwargs (Any, default: {} ) –

    Additional keyword arguments forwarded to :func:seaborn.scatterplot. They override the defaults computed here, so anything seaborn understands (palette, style, size, markers, ...) can be tuned from the call site.

Property Plots

VarNumberBarChartPlot

Bases: BarPlot

Bar chart showing the number of variables per model.

Every display option is inherited and can be set when the plot is constructed, e.g. VarNumberBarChartPlot(annotate=False, file_formats=("pgf", "png")). BarPlot documents the colours, error bars, value annotations and grouping by a feature; SeabornPlot the figure size and the output formats.

Attributes:

  • figure_filename (str) –

    Stem of the written figure files, by default "var_number_bar_chart".

Examples:

>>> bench.add_feature(name="var_count", feature=VarNumberFeature())
>>> bench.add_plot(name="var_number", plot=VarNumberBarChartPlot())

theme: Theme | None = Theme() class-attribute instance-attribute

The seaborn theme the figure is drawn under, and the gridlines behind the marks.

A figure is read against its axis, so it is drawn with the lines that make that possible unless asked otherwise. None takes the theme and the grid away.

missing: Missing = Missing() class-attribute instance-attribute

What becomes of the values the plot cannot draw, and how it says they were there.

option_bundles: dict[str, type[OptionBundle]] = {'style': PlotStyle} class-attribute

Constructor arguments that configure several fields at once, and the bundle each takes.

A style is spread over the bundle fields rather than stored, so a benchmark can hand the same look to every plot while each keeps what it says itself. Applied least specific first: the shared style, then a bundle passed to the plot, then a flat option.

aggregation: Aggregation = Aggregation.MEAN class-attribute instance-attribute

Aggregation applied to the values of an x category, by default their mean.

errorbars: ErrorBars | None = ErrorBars() class-attribute instance-attribute

The error bars drawn on top of the bars: what they show, their colour and caps.

None draws none, the same as ErrorBars(spec=None).

annotation: Annotation | None = None class-attribute instance-attribute

The values written above the bars - how they are formatted and how large.

None, the default, writes none: a bar chart is read off its axis, and a number above every bar is worth its clutter only when the exact value is the point. Pass an Annotation to turn them on, empty for the defaults.

grouping: Dimension | None = None class-attribute instance-attribute

What splits each bar into a group of bars.

One of the groupers - ModelDimension, AlgorithmDimension, FeatureDimension, ParameterDimension - or None, which leaves the bars ungrouped.

run(benchmark_results: BenchmarkResultContainer, save_dir: str | None = None) -> None

Generate plot output from benchmark results.

Parameters:

draw_into(axes: Axes, benchmark_results: BenchmarkResultContainer) -> None

Run this plot onto an existing axes rather than into a figure of its own.

Everything the plot draws goes through the pyplot state, so making that axes the current one is enough to redirect it. The figure, the files, and the window stay the caller's business - which is what lets several plots share one figure, e.g. the summary grid.

Parameters:

  • axes (Axes) –

    The axes to draw on.

  • benchmark_results (BenchmarkResultContainer) –

    Aggregated benchmark data handed to :meth:run.

resolve_missing(df: pd.DataFrame, column: str, *, by: str | None = None, within: str | None = None) -> tuple[pd.DataFrame, dict[tuple[str, str], int]]

Return the drawable rows and how many of them were not, per category.

A metric with nothing to report says so with a None or an infinity - a time to solution of a run that never reached the optimum is the usual one, since the expected time to something that did not happen is unbounded. Neither is a height a bar can have, and leaving them in poisons the aggregate: one infinity turns the mean of an algorithm into an infinity, and a missing value silently shortens it.

What happens to them is :attr:missing, and by default it is nothing: the plot raises rather than quietly showing a mean over fewer models than it claims. Asked to carry on, it either leaves them out or fills them from the values that could be drawn - Missing(policy="max") puts them just past the tallest bar, which is where "worse than everything here" belongs. Either way they are counted and a warning is logged: a bar resting on half its models is a different statement from one resting on all of them, and that is not visible in the bar itself.

Parameters:

  • df (DataFrame) –

    The plotting data.

  • column (str) –

    Column holding the plotted value.

  • by (str | None, default: None ) –

    Column whose categories the missing values are counted per, e.g. the x-axis of a bar plot. Without one they are counted under "".

  • within (str | None, default: None ) –

    Column that splits those categories further, e.g. the grouping of a bar plot. Counting per group is what lets a figure mark the one bar of a group that lost values rather than the whole category it sits in.

Returns:

  • tuple[DataFrame, dict[tuple[str, str], int]] –

    The rows to draw, and the number of missing values per category and group - the group is "" where there is none. The mapping is empty when nothing was missing.

Raises:

  • PlotMissingValuesError –

    If values are missing and the policy is "raise".

place_legend(axes: Axes, handles: list[Any] | None = None, labels: list[str] | None = None) -> None

Put the legend beside the axes, whatever drew it.

Outside the axes for a figure of its own: a legend inside sits on top of the data, and which corner is free depends on the run rather than on the plot - the figure would move its own key around as the numbers change. Beside it, the key is always in the same place and covers nothing.

A panel of someone else's figure is the exception. The room beside it belongs to the panel next to it, so a key anchored there is drawn over a neighbour rather than over the data; inside the panel it stays within the space the plot was given.

Parameters:

  • axes (Axes) –

    The axes the plot was drawn on.

  • handles (list[Any] | None, default: None ) –

    Legend handles, by default the ones already on the axes.

  • labels (list[str] | None, default: None ) –

    Their labels, by default the ones already on the axes.

note_missing(handles: list[Any], labels: list[str], missing: dict[tuple[str, str], int]) -> None

Add the legend entry that says how many values the figure could not draw.

What a plot can say beyond that depends on what it draws. A bar has a slot of its own to put a cross under, so BarPlot marks the categories themselves; a point in a cloud or a step of a sweep has no slot, and the count in the key is the whole statement there - enough that a filled value is not read as a measured one.

Parameters:

  • handles (list[Any]) –

    Legend handles, extended in place.

  • labels (list[str]) –

    Their labels, extended in place alongside handles.

  • missing (dict[tuple[str, str], int]) –

    Number of missing values per category and group, as counted by :meth:resolve_missing.

apply_theme() -> None

Install the seaborn theme this plot is drawn under, unless it has none.

The theme is matplotlib's global state rather than a property of one figure, so it is installed before the figure is built and left in place afterwards: a benchmark themes its plots by handing every one of them the same Theme, not by each plot putting the previous look back.

apply_grid(axes: Axes) -> None

Draw the gridlines the theme asks for, behind everything else on axes.

Parameters:

  • axes (Axes) –

    The axes the plot was drawn on.

setup_figure() -> None

Create a matplotlib figure, unless the plot is drawing into a shared axes.

save_figure(save_dir: str) -> list[Path]

Write the current figure to save_dir once per configured file format.

Parameters:

  • save_dir (str) –

    Directory to save the figure into. Created if it does not exist.

Returns:

  • list[Path] –

    Paths that were written successfully.

finalize_plot(xlabel: str, ylabel: str, title: str, ylim: tuple[float, float] | None = None, x_rotation: int = 45, save_dir: str | None = None) -> None

Apply common axis labels, title, limits, and display behavior.

Parameters:

  • xlabel (str) –

    Label for the x-axis.

  • ylabel (str) –

    Label for the y-axis.

  • title (str) –

    Plot title.

  • ylim (tuple[float, float] | None, default: None ) –

    Lower and upper y-axis limits, by default None.

  • x_rotation (int, default: 45 ) –

    Rotation angle for x-axis tick labels, by default 45.

  • save_dir (str | None, default: None ) –

    Directory to save the figure into, by default None.

apply_grouping(benchmark_results: BenchmarkResultContainer, rows: list[dict[str, Any]]) -> dict[str, Any]

Split rows into groups along :attr:grouping.

What that means is the grouper's business - a column of the plotted data, a value looked up per model, or a setting the algorithms were configured with - and so is deciding that it does not apply, in which case the bars stay ungrouped.

Parameters:

  • benchmark_results (BenchmarkResultContainer) –

    Benchmark data the feature results and algorithm configurations are read from.

  • rows (list[dict[str, Any]]) –

    Row-oriented plot data, annotated - and, where a grouping applies to only part of the data, reduced - in place.

Returns:

  • dict[str, Any] –

    Keyword arguments to forward to :meth:create. Empty when no grouping applies, so call sites can splat it unconditionally.

draw(*, benchmark_results: BenchmarkResultContainer, rows: list[dict[str, Any]], save_dir: str | None = None, **overrides: Any) -> None

Group rows and draw them with the display configuration of this plot.

This is what turns the declared fields - :attr:x, :attr:title, :attr:hline and the rest - into a :meth:create call, so a subclass only has to say which rows it plots. Doing it in one place is also what keeps :attr:group_by working for every bar plot rather than for those that remember to apply it.

Parameters:

  • benchmark_results (BenchmarkResultContainer) –

    Benchmark data, used to look up the groups of a feature :attr:group_by.

  • rows (list[dict[str, Any]]) –

    Row-oriented plot data.

  • save_dir (str | None, default: None ) –

    Directory to save the figure into, by default None.

  • **overrides (Any, default: {} ) –

    Keyword arguments forwarded to :meth:create, overriding the fields.

transform_rows(rows: list[dict[str, Any]], x: str | None, group: str | None) -> list[dict[str, Any]]

Return the rows to plot, by default the rows as they are.

A subclass that has to reduce its rows before they are drawn - pooling counts into a single ratio, say - overrides this rather than :meth:run, so it keeps the shared grouping and display handling. It is told what the bars and the groups turned out to be, since that is what a row has to keep to stay one of them.

Parameters:

  • rows (list[dict[str, Any]]) –

    Row-oriented plot data, already annotated with the dimensions' columns.

  • x (str | None) –

    Column the bars are drawn per, or None when the plot has no rows.

  • group (str | None) –

    Column the bars are split by, or None when they are ungrouped.

Returns:

create(*, rows: list[dict[str, Any]], xlabel: str, ylabel: str, title: str, x: str = 'x', y: str = 'y', aggregation: Aggregation = Aggregation.MEAN, errorbar: ErrorBar | str = AUTO_ERRORBAR, hue: str | None = None, hline: float | None = None, hline_label: str | None = None, hcolor: str = REFERENCE_LINE_COLOUR, baseline: float | None = None, ylim: tuple[float, float] | None = None, legend: bool = False, save_dir: str | None = None, **kwargs: Any) -> None

Create a bar plot from row-oriented data.

Parameters:

  • rows (dict[str, Any]) –

    Row-oriented mapping used to construct the plotting DataFrame.

  • xlabel (str) –

    Label for the x-axis.

  • ylabel (str) –

    Label for the y-axis.

  • title (str) –

    Plot title.

  • x (str, default: 'x' ) –

    Column name mapped to the x-axis, by default "x".

  • y (str, default: 'y' ) –

    Column name mapped to the y-axis, by default "y".

  • aggregation (Aggregation, default: MEAN ) –

    Aggregation strategy applied by seaborn, by default Aggregation.MEAN.

  • errorbar (ErrorBar | str, default: AUTO_ERRORBAR ) –

    Seaborn error bar specification ("sd", ("ci", 95), None to disable). By default "auto", which takes the error bar from aggregation: the spread of the samples for means, none for min/max.

  • hue (str | None, default: None ) –

    Optional grouping column for grouped bars, by default None.

  • hline (float | None, default: None ) –

    Optional horizontal reference line value, by default None.

  • hline_label (str | None, default: None ) –

    Legend label for the horizontal reference line, by default None.

  • hcolor (str, default: REFERENCE_LINE_COLOUR ) –

    Colour of the horizontal reference line, by default black.

  • baseline (float | None, default: None ) –

    Height of a solid black baseline marking where the bars start, by default None. Unlike hline it carries no label and stays out of the legend - it says where zero is, it does not name a target.

  • ylim (tuple[float, float] | None, default: None ) –

    Lower and upper y-axis limits, by default None.

  • legend (bool, default: False ) –

    Whether seaborn should create a legend for hue groups, by default False.

  • save_dir (str | None, default: None ) –

    Directory to save the figure into, by default None.

  • **kwargs (Any, default: {} ) –

    Additional keyword arguments forwarded to :func:seaborn.barplot. They override the defaults computed here, so anything seaborn understands (palette, saturation, capsize, err_kws, ...) can be tuned from the call site.

annotation_text(value: float) -> str

Return the text written above a bar of value.

Applies :attr:annotate_format, unless :attr:annotate_max_decimals allows the value to be written as a plain decimal instead of in scientific notation. A value read off a percent axis is written as a percentage, so the annotation says the same thing as the axis it stands on - unless a format was asked for, which wins.

Parameters:

  • value (float) –

    The aggregated value of one bar.

Returns:

  • str –

    The annotation, e.g. "0.000057" rather than "5.67e-05".