Validators#

Outbound and inbound validation pipelines for data transformation.

Reusable validator functions and composable pipelines for value normalization.

Provides functions for composable inbound/outbound value normalization (e.g., converting between pandas DataFrames and dictionaries).

ionworks.validators.set_dataframe_backend(backend)[source]#

Set the default DataFrame backend for data fetching.

This overrides the IONWORKS_DATAFRAME_BACKEND environment variable.

Parameters:

backend (str) – DataFrame backend to use: “polars” or “pandas”.

Raises:

ValueError – If backend is not “polars” or “pandas”.

Return type:

None

ionworks.validators.get_dataframe_backend()[source]#

Get the current DataFrame backend setting.

Returns:

Current backend: “polars” or “pandas”.

Return type:

str

ionworks.validators.CAPACITY_ENERGY_INTEGRAL_TOLERANCE: float = 0.1#

Default relative-error tolerance for capacity/energy integral checks. Used by validate_capacity_energy_from_current_power() and by ionworksdata.transform.fix_swapped_charge_discharge_columns so both agree on the same definition of “within tolerance”.

class ionworks.validators.IssueCode(*values)[source]#

Bases: StrEnum

Stable identifiers for measurement validation findings.

Use these — not substrings of the human-readable message — when branching on issue identity. Members are str-compatible.

CURRENT_SIGN_REVERSED = 'current_sign_reversed'#
CURRENT_SIGN_INDETERMINATE = 'current_sign_indeterminate'#
CURRENT_SIGN_UNSIGNED = 'current_sign_unsigned'#
CUMULATIVE_VALUE_NOT_RESET = 'cumulative_value_not_reset'#
CUMULATIVE_VALUE_DECREASED = 'cumulative_value_decreased'#
STEP_TOO_FEW_POINTS = 'step_too_few_points'#
TIME_DOES_NOT_START_AT_ZERO = 'time_does_not_start_at_zero'#
TIME_NOT_MONOTONIC = 'time_not_monotonic'#
STEP_COUNT_MISSING = 'step_count_missing'#
STEP_COUNT_DOES_NOT_START_AT_ZERO = 'step_count_does_not_start_at_zero'#
STEP_COUNT_NON_SEQUENTIAL = 'step_count_non_sequential'#
CYCLE_CHANGES_WITHIN_STEP = 'cycle_changes_within_step'#
OCP_VOLTAGE_COLUMN_MISSING = 'ocp_voltage_column_missing'#
OCP_X_AXIS_COLUMN_MISSING = 'ocp_x_axis_column_missing'#
TIME_SERIES_ROW_COUNT_EXCEEDED = 'time_series_row_count_exceeded'#
TIME_GAP_TOO_LARGE = 'time_gap_too_large'#
VOLTAGE_CONTINUITY = 'voltage_continuity'#
CONSECUTIVE_SAME_DIRECTION_FULL_STEPS = 'consecutive_same_direction_full_steps'#
STEP_CAPACITY_EXCEEDS_RATED = 'step_capacity_exceeds_rated'#
CHARGE_DISCHARGE_CAPACITY_COLUMNS_SWAPPED = 'charge_discharge_capacity_columns_swapped'#
CHARGE_DISCHARGE_ENERGY_COLUMNS_SWAPPED = 'charge_discharge_energy_columns_swapped'#
DISCHARGE_CAPACITY_INTEGRAL_MISMATCH = 'discharge_capacity_integral_mismatch'#
CHARGE_CAPACITY_INTEGRAL_MISMATCH = 'charge_capacity_integral_mismatch'#
DISCHARGE_ENERGY_INTEGRAL_MISMATCH = 'discharge_energy_integral_mismatch'#
CHARGE_ENERGY_INTEGRAL_MISMATCH = 'charge_energy_integral_mismatch'#
EIS_COLUMNS_MISSING = 'eis_columns_missing'#
EIS_ZIM_SIGN_REVERSED = 'eis_zim_sign_reversed'#
EIS_IMPEDANCE_MAGNITUDE_IMPLAUSIBLE = 'eis_impedance_magnitude_implausible'#
START_TIME_IN_FUTURE = 'start_time_in_future'#
END_TIME_IN_FUTURE = 'end_time_in_future'#
END_TIME_BEFORE_START_TIME = 'end_time_before_start_time'#
DURATION_MISMATCH = 'duration_mismatch'#
class ionworks.validators.ValidationIssue(code, message, severity='error', payload=<factory>)[source]#

Bases: object

Structured measurement validation finding.

Returned by the validate_* functions and carried by MeasurementValidationError. Branch on code rather than parsing message; messages are for display only and may be reworded between releases.

Parameters:
  • code (IssueCode) – Stable identifier of the finding.

  • message (str) – Human-readable description of the failure, including any fix hint.

  • severity ({"error", "warning"}, optional) – "error" (the default) causes validate_measurement_data() to raise. "warning" is emitted via warnings.warn() and is not attached to the resulting MeasurementValidationError — callers who need programmatic access to warning-severity issues must invoke the underlying validate_* producer directly rather than relying on e.has_code(...).

  • payload (dict, optional) – Structured details: step indices, column names, observed values, thresholds. Keys depend on code.

code: IssueCode#
message: str#
severity: Literal['error', 'warning'] = 'error'#
payload: dict[str, Any]#
__init__(code, message, severity='error', payload=<factory>)#
Parameters:
Return type:

None

exception ionworks.validators.MeasurementValidationError(message, errors=None)[source]#

Bases: IonworksError

Exception raised when measurement data validation fails.

errors is a list of structured ValidationIssue records. Branch on issue.code rather than parsing issue.message.

Parameters:
Return type:

None

__init__(message, errors=None)[source]#

Initialize the IonworksError.

Parameters:
  • message (str | dict[str, Any]) – Error message string or dict containing error details. Supports both the legacy {"detail": ...} format and the new standardized {"error_code": ..., "message": ..., "detail": ...} format.

  • status_code (int | None) – Optional HTTP status code.

  • errors (list[ValidationIssue] | None)

Return type:

None

errors: list[ValidationIssue]#
has_code(code)[source]#

Return True if any issue matches code.

Parameters:

code (IssueCode)

Return type:

bool

ionworks.validators.positive_current_is_charge(t, current, voltage)[source]#

Determine whether positive current corresponds to charging.

Fits an OCV-R equivalent-circuit model V = OCV(SOC) - I * R0 under two sign-convention hypotheses (positive = discharge vs. positive = charge). OCV is linear in SOC, constrained to a non-negative slope (monotonically increasing); R0 is a single positive scalar. When the sign convention is wrong the SOC axis is inverted, forcing the OCV slope toward zero. The convention producing the larger OCV slope is selected.

Parameters:
  • t (np.ndarray) – Time values [s].

  • current (np.ndarray) – Current values [A].

  • voltage (np.ndarray) – Voltage values [V].

Returns:

  • is_charge (bool) – True if positive current is charging, False if discharging. Returns False when there is insufficient data.

  • p_value (float) – Confidence metric in [0, 1]. Lower values indicate higher confidence. Computed as the sum of two ratios clipped to [0, 1]: the ratio of the smaller to the larger OCV slope (delta_ratio) and the ratio of the winner’s SSE to the loser’s SSE (sse_ratio). Returns 1.0 when the result is ambiguous or data is insufficient.

Return type:

tuple[bool, float]

ionworks.validators.validate_positive_current_is_discharge(df, current_col='Current [A]', voltage_col='Voltage [V]', time_col='Time [s]', step_col=None, rest_tol=0.001, relative_rest_tol_frac=0.02, relative_rest_tol_percentile=95.0)[source]#

Validate that positive current corresponds to discharge.

Discharge should cause voltage to decrease. This function analyzes the relationship between current direction and voltage change to verify the sign convention is correct.

Fits an OCV-R ECM per step, then uses a confidence vote across steps weighted by the trapezoidal integral of |I(t)| dt over each step so that steps actually moving charge dominate the decision and long near-zero-current voltage holds contribute negligible weight.

Parameters:
  • df (DataFrame) – Time series data with current and voltage columns (pandas or polars).

  • current_col (str) – Name of the current column.

  • voltage_col (str) – Name of the voltage column.

  • time_col (str) – Name of the time column.

  • step_col (str, optional) – Name of the step column. If provided, analyzes per-step. Otherwise, infers steps from current sign changes.

  • rest_tol (float) – Tolerance for considering current as zero (rest).

  • relative_rest_tol_frac (float) – Fraction of a robust current scale used to set a relative rest threshold. The robust scale is the percentile of abs(current) given by relative_rest_tol_percentile.

  • relative_rest_tol_percentile (float) – Percentile in [0, 100] used to estimate the robust current scale for the relative threshold. The effective rest threshold is min(rest_tol, relative_rest_tol) so the stricter (smaller) threshold takes precedence.

Returns:

List of validation issues. Empty if validation passes.

Return type:

list[ValidationIssue]

ionworks.validators.validate_cumulative_values_reset_per_step(df, step_col='Step count', cumulative_cols=None, tolerance=1e-06)[source]#

Validate cumulative values reset to ~0 at each step and only increase.

Parameters:
  • df (DataFrame) – Time series data (pandas or polars).

  • step_col (str) – Name of the column containing step numbers.

  • cumulative_cols (list[str], optional) – List of cumulative column names to validate. If None, checks for common capacity and energy columns.

  • tolerance (float) – Tolerance for considering a value as “zero” at step start.

Returns:

List of validation issues. Empty if validation passes.

Return type:

list[ValidationIssue]

ionworks.validators.validate_minimum_points_per_step(df, step_col='Step count', min_points=2)[source]#

Validate that each step has at least a minimum number of data points.

Parameters:
  • df (DataFrame) – Time series data (pandas or polars).

  • step_col (str) – Name of the column containing step numbers.

  • min_points (int) – Minimum number of points required per step.

Returns:

List of validation issues. Empty if validation passes.

Return type:

list[ValidationIssue]

ionworks.validators.validate_time_starts_at_zero(df, tolerance=1e-06)[source]#

Validate that ‘Time [s]’ starts at 0.

Parameters:
  • df (DataFrame) – Time series data (pandas or polars).

  • tolerance (float) – Tolerance for considering the start value as zero.

Returns:

List of validation issues. Empty if validation passes.

Return type:

list[ValidationIssue]

ionworks.validators.validate_time_monotonic(df, time_col='Time [s]', tolerance=1e-12)[source]#

Validate that the time column is monotonically non-decreasing.

Parameters:
  • df (DataFrame) – Time series data (pandas or polars).

  • time_col (str) – Name of the time column.

  • tolerance (float) – Numerical tolerance; time[i] must be >= time[i-1] - tolerance.

Returns:

List of validation issues. Empty if validation passes.

Return type:

list[ValidationIssue]

ionworks.validators.validate_step_count_sequential(df)[source]#

Validate that ‘Step count’ exists, starts at 0, and increases by 1.

Parameters:

df (DataFrame) – Time series data (pandas or polars).

Returns:

List of validation issues. Empty if validation passes.

Return type:

list[ValidationIssue]

ionworks.validators.validate_cycle_constant_within_step(df, step_col='Step count', cycle_col=None)[source]#

Validate that cycle number does not change within a step.

Parameters:
  • df (DataFrame) – Time series data (pandas or polars).

  • step_col (str) – Name of the column containing step numbers.

  • cycle_col (str, optional) – Name of the column containing cycle numbers. If None, tries common names.

Returns:

List of validation issues. Empty if validation passes.

Return type:

list[ValidationIssue]

ionworks.validators.validate_ocp_columns(df)[source]#

Validate that OCP data has required columns.

Checks that the DataFrame contains: 1. A ‘Voltage [V]’ column 2. At least one x-axis column: ‘Capacity [A.h]’, ‘Stoichiometry’, or ‘SOC’

Parameters:

df (DataFrame) – Time series data (pandas or polars).

Returns:

List of validation issues. Empty if validation passes.

Return type:

list[ValidationIssue]

ionworks.validators.EIS_FREQ_COL = 'Frequency [Hz]'#

Default EIS column names (BioLogic/Gamry reader convention).

ionworks.validators.EIS_RESISTANCE_CAPACITY_PRODUCT: float = 0.05#

Roughly cell-size-invariant ohmic resistance-capacity product [Ohm.A.h], used to estimate a plausible ohmic-resistance scale from rated capacity alone (R0 ~ product / capacity, since larger cells have proportionally more electrode area in parallel). For Li-ion this clusters around 0.05-0.09 Ohm.A.h (power cells lower, small high-energy cells higher; ~1 order of magnitude spread), so 0.05 is a central-but-conservative anchor. It is only an order-of-magnitude reference for the deliberately wide (default 100x) EIS magnitude sanity band, so the exact value is not load-bearing.

ionworks.validators.validate_eis_columns(df, freq_col='Frequency [Hz]', zre_col='Z_Re [Ohm]', zim_col='Z_Im [Ohm]')[source]#

Validate that EIS data has the required frequency and impedance columns.

Parameters:
  • df (DataFrame) – EIS spectrum data (pandas or polars).

  • freq_col (str, optional) – Name of the frequency column. Defaults to "Frequency [Hz]".

  • zre_col (str, optional) – Name of the real-impedance column. Defaults to "Z_Re [Ohm]".

  • zim_col (str, optional) – Name of the imaginary-impedance column. Defaults to "Z_Im [Ohm]".

Returns:

A single EIS_COLUMNS_MISSING issue listing the missing and available columns, or an empty list when all required columns are present.

Return type:

list[ValidationIssue]

ionworks.validators.validate_eis_zim_sign(df, freq_col='Frequency [Hz]', zre_col='Z_Re [Ohm]', zim_col='Z_Im [Ohm]', reversed_fraction_threshold=0.75)[source]#

Validate that the EIS Z_Im sign convention is not reversed.

The database stores the raw imaginary part Z_Im = Im(Z), which is negative across the capacitive arc; only a possible high-frequency inductive loop is positive. This check sorts by frequency, trims the inductive tail (points above the R0 = min(Z_Re) intercept), and flags the spectrum when the capacitive band is predominantly positive — i.e. the whole spectrum has been stored with a flipped sign. Points are weighted by |Z_Im| so the near-zero points around the real-axis crossing do not dominate the vote.

Parameters:
  • df (DataFrame) – EIS spectrum data (pandas or polars).

  • freq_col (str, optional) – Name of the frequency column. Defaults to "Frequency [Hz]".

  • zre_col (str, optional) – Name of the real-impedance column. Defaults to "Z_Re [Ohm]".

  • zim_col (str, optional) – Name of the imaginary-impedance column. Defaults to "Z_Im [Ohm]".

  • reversed_fraction_threshold (float, optional) – Minimum |Z_Im|-weighted fraction of the capacitive band that must be positive to flag a reversed sign. Defaults to 0.75 (75 %).

Returns:

A single EIS_ZIM_SIGN_REVERSED issue when the sign appears flipped, or an empty list otherwise (including when columns are missing or there are too few points).

Return type:

list[ValidationIssue]

ionworks.validators.validate_eis_impedance_magnitude(df, rated_capacity, factor=100.0, r_times_q=0.05, zre_col='Z_Re [Ohm]')[source]#

Check that the EIS ohmic resistance is within a wide plausibility band.

Estimates a plausible ohmic-resistance scale from the rated capacity alone via the roughly cell-size-invariant product R0 ~ r_times_q / rated_capacity, then flags the spectrum only when the measured intercept R0 = min(Z_Re) is more than factor times away from it in either direction. The band is deliberately wide (default 100x) so it fires only on order-of-magnitude errors — e.g. an impedance stored in the wrong unit — and never on a plausible spectrum. The exact r_times_q value is unimportant at this tolerance.

Parameters:
  • df (DataFrame) – EIS spectrum data (pandas or polars).

  • rated_capacity (float | None) – Rated (nominal) cell capacity in A.h. The check is skipped when this is None or non-positive.

  • factor (float, optional) – Half-width of the sanity band as a multiplicative factor around the expected ohmic resistance. Defaults to 100.0.

  • r_times_q (float, optional) – Order-of-magnitude resistance-capacity product [Ohm.A.h] used to estimate the expected ohmic resistance. Defaults to EIS_RESISTANCE_CAPACITY_PRODUCT.

  • zre_col (str, optional) – Name of the real-impedance column. Defaults to "Z_Re [Ohm]".

Returns:

A single EIS_IMPEDANCE_MAGNITUDE_IMPLAUSIBLE issue when R0 is outside the band, or an empty list otherwise (including when the rated capacity is unavailable, the column is missing, or R0 is non-positive).

Return type:

list[ValidationIssue]

ionworks.validators.validate_time_series_row_count(df, max_rows=1000)[source]#

Validate that the time series does not exceed the maximum row count.

Datasets larger than max_rows should be uploaded via the standard upload flow and then referenced with "db:<measurement_id>" or iwdata.DataLoader.from_db(MEASUREMENT_ID) in pipeline configurations.

Parameters:
  • df (DataFrame) – Time series data (pandas or polars).

  • max_rows (int) – Maximum allowed number of rows.

Returns:

List of validation issues. Empty if validation passes.

Return type:

list[ValidationIssue]

ionworks.validators.validate_time_gaps(df, time_col='Time [s]', max_gap_seconds=18000)[source]#

Validate that there are no large gaps between consecutive time samples.

A gap longer than max_gap_seconds between two consecutive rows is almost always a sign that a chunk of cycling was dropped — for example, a rest period was recorded as elapsed time in the file but the intermediate rows were stripped, leaving the capacity integral to count current across the unrecorded interval.

Parameters:
  • df (DataFrame) – Time series data with a time column (pandas or polars).

  • time_col (str, optional) – Name of the time column. Defaults to "Time [s]".

  • max_gap_seconds (float, optional) – Maximum allowed gap between consecutive time samples, in seconds. Defaults to 5 * 3600 (5 hours).

Returns:

List of validation issues. Empty if validation passes.

Return type:

list[ValidationIssue]

ionworks.validators.validate_voltage_continuity(df, voltage_window, voltage_col='Voltage [V]', jump_fraction=0.8, max_bad_fraction=0.05)[source]#

Validate that consecutive rows do not exhibit unphysical voltage jumps.

Computes the absolute voltage difference between every pair of consecutive rows and counts the fraction that exceed jump_fraction of the rated voltage window (V_max - V_min). If more than max_bad_fraction of row pairs exceed that threshold, the data is flagged as likely being out of chronological order (for example, after a faulty partition_by("Cycle_raw") grouping that interleaves pulse and rest rows from different cycles).

Parameters:
  • df (DataFrame) – Time series data with a voltage column (pandas or polars).

  • voltage_window (tuple[float, float]) – (V_min, V_max) rated voltage window of the cell, typically the lower/upper cutoff voltages. Only the span V_max - V_min is used.

  • voltage_col (str, optional) – Name of the voltage column. Defaults to "Voltage [V]".

  • jump_fraction (float, optional) – Fraction of the voltage window considered the maximum physically plausible single-row voltage change. Defaults to 0.80 (80 %).

  • max_bad_fraction (float, optional) – Maximum fraction of consecutive row pairs allowed to exceed the jump threshold before the check fails. Defaults to 0.05 (5 %).

Returns:

List of validation issues. Empty if validation passes.

Return type:

list[ValidationIssue]

ionworks.validators.validate_consecutive_same_direction_full_steps(steps_df, rated_capacity, step_type_col='Step type', discharge_capacity_col='Discharge capacity [A.h]', charge_capacity_col='Charge capacity [A.h]', step_count_col='Step count', full_step_fraction=2.0)[source]#

Validate that no two consecutive full-capacity steps share the same direction.

Walks the step summary in order. For each constant-current step (identified by its Step type) that delivers more than full_step_fraction of the rated capacity, tracks the direction. If two such steps appear consecutively with the same direction (ignoring rest / EIS / unknown steps, which do not reset the streak), the check fails. This catches measurements that bundle multiple independent experiments — e.g. a discharge-rate file concatenating several CC discharges into a single measurement.

Parameters:
  • steps_df (DataFrame) – Step summary dataframe (as produced by ionworksdata.steps.summarize).

  • rated_capacity (float | None) – Rated (nominal) cell capacity in A.h. Used together with full_step_fraction to decide whether a step counts as a full charge / discharge. The check is skipped when this is None or non-positive.

  • step_type_col (str, optional) – Name of the column containing the step type labels ("Rest", "Constant current discharge", etc.). Defaults to "Step type".

  • discharge_capacity_col (str, optional) – Name of the per-step discharge capacity column.

  • charge_capacity_col (str, optional) – Name of the per-step charge capacity column.

  • step_count_col (str, optional) – Name of the per-step identifier column, reported in error messages.

  • full_step_fraction (float, optional) – Multiple of rated capacity above which a step is considered a full (and then some) charge / discharge. Defaults to 2.0 — i.e. only steps delivering more than 2× rated capacity trigger the check, which tolerates one or two full cycles being captured in a single step while still catching runaway concatenations.

Returns:

List of validation issues. Empty if validation passes.

Return type:

list[ValidationIssue]

ionworks.validators.validate_step_capacity_within_rated(steps_df, rated_capacity, discharge_capacity_col='Discharge capacity [A.h]', charge_capacity_col='Charge capacity [A.h]', step_count_col='Step count', max_ratio=5.0)[source]#

Soft-check that no single step exceeds max_ratio × rated capacity.

A single step accumulating more capacity than several full charges or discharges of the cell usually indicates wrong step boundaries or that the capacity integral was inflated across unrecorded time gaps. Returns a list of warnings; callers may emit them via warnings.warn() instead of raising.

Parameters:
  • steps_df (DataFrame) – Step summary dataframe (as produced by ionworksdata.steps.summarize).

  • rated_capacity (float | None) – Rated (nominal) cell capacity in A.h. The check is skipped when this is None or non-positive.

  • discharge_capacity_col (str, optional) – Name of the per-step discharge capacity column.

  • charge_capacity_col (str, optional) – Name of the per-step charge capacity column.

  • step_count_col (str, optional) – Name of the per-step identifier column, reported in warning messages.

  • max_ratio (float, optional) – Maximum allowed ratio of per-step capacity to rated capacity. Defaults to 5.0 (500 %).

Returns:

List of validation issues. Empty if no step exceeds the threshold.

Return type:

list[ValidationIssue]

ionworks.validators.validate_charge_discharge_column_direction(steps_df, mean_current_col='Mean current [A]', min_current_col='Min current [A]', max_current_col='Max current [A]', std_current_col='Std current [A]', duration_col='Duration [s]', step_count_col='Step count', discharge_capacity_col='Discharge capacity [A.h]', charge_capacity_col='Charge capacity [A.h]', discharge_energy_col='Discharge energy [W.h]', charge_energy_col='Charge energy [W.h]', rest_tol=0.001, min_duration_s=60.0, max_std_to_mean_ratio=0.5, dominance_ratio=4.0, swapped_fraction_threshold=0.75)[source]#

Detect swapped charge/discharge capacity (and energy) columns in a step summary.

With the platform sign convention (positive current = discharge), a step that only discharges must accumulate its A.h in the discharge column, and a step that only charges must accumulate in the charge column. A cycler export with the two column labels inverted is symmetric under every cumulative-reset and magnitude check, so this is the only signal that catches it: for each step whose current direction is unambiguous, compare the sign of the mean current against which column of the pair actually accumulated, and flag the pair when the clear majority of such steps put their capacity in the opposite-direction column.

A step votes only when the evidence is unambiguous on every axis:

  • its mean current is clearly non-zero (|mean| >= rest_tol),

  • it lasted at least min_duration_s (when a duration column is available),

  • its current did not change sign — the min/max currents stay on the same side of zero (within rest_tol), or the standard deviation is small relative to |mean|; either piece of evidence qualifies, so a single opposite-sign boundary sample inherited from the preceding step does not silence an otherwise one-directional step. Bidirectional (drive-cycle) steps legitimately accumulate in both columns and are excluded by this test, and

  • one column of the pair clearly dominates the other (dominance_ratio), so a step that accumulated comparably in both columns never votes.

Votes are weighted by the capacity (or energy) the step moved, so a handful of tiny glitch steps cannot outvote real cycling. The capacity pair and the energy pair are evaluated independently; energy direction is also judged from the mean current, since terminal voltage is positive and power therefore shares the current’s sign.

Parameters:
  • steps_df (DataFrame) – Step summary dataframe (as produced by ionworksdata.steps.identify), pandas or polars.

  • mean_current_col (str, optional) – Name of the per-step mean current column [A]. The check is skipped entirely when this column is absent.

  • min_current_col (str, optional) – Names of the per-step min/max current columns [A], used to establish that a step’s current never changed sign.

  • max_current_col (str, optional) – Names of the per-step min/max current columns [A], used to establish that a step’s current never changed sign.

  • std_current_col (str, optional) – Name of the per-step current standard deviation column [A]. Alternative sign-unambiguity evidence: the step also qualifies when std <= max_std_to_mean_ratio * |mean|, even when its min/max fail the same-sign test (e.g. polluted by a boundary sample).

  • duration_col (str, optional) – Name of the per-step duration column [s]. Steps shorter than min_duration_s (or with NaN duration) do not vote. No duration filter is applied when the column is absent.

  • step_count_col (str, optional) – Name of the per-step identifier column, reported in the issue payload.

  • discharge_capacity_col (str, optional) – Names of the per-step capacity pair [A.h]. The pair is skipped when either column is absent.

  • charge_capacity_col (str, optional) – Names of the per-step capacity pair [A.h]. The pair is skipped when either column is absent.

  • discharge_energy_col (str, optional) – Names of the per-step energy pair [W.h]. The pair is skipped when either column is absent.

  • charge_energy_col (str, optional) – Names of the per-step energy pair [W.h]. The pair is skipped when either column is absent.

  • rest_tol (float, optional) – Current magnitude [A] below which a step counts as rest and does not vote. Also the tolerance for the min/max same-sign test, so a discharge step whose current briefly touches zero still qualifies. Defaults to 1e-3.

  • min_duration_s (float, optional) – Minimum step duration [s] for a step to vote. Defaults to 60.

  • max_std_to_mean_ratio (float, optional) – Maximum std / |mean| for a step to qualify via the standard-deviation fallback. Defaults to 0.5.

  • dominance_ratio (float, optional) – Minimum ratio of the larger to the smaller column value for the step’s accumulation to count as clearly one-directional. Defaults to 4.0.

  • swapped_fraction_threshold (float, optional) – Minimum weighted fraction of voting steps that must accumulate in the wrong column to flag the pair as swapped. Defaults to 0.75.

Returns:

At most one issue per pair (CHARGE_DISCHARGE_CAPACITY_COLUMNS_SWAPPED and/or CHARGE_DISCHARGE_ENERGY_COLUMNS_SWAPPED). Empty when neither pair appears swapped or there is not enough unambiguous evidence.

Return type:

list[ValidationIssue]

ionworks.validators.step_boundary_idx(step_data)[source]#

Per-row index of the most recent step start.

Subtracting series[step_boundary_idx(step_data)] from a running cumulative series resets it to 0 at every change in step_data. Compute this once and pass it into multiple running_step_reset_integral() calls when they share the same step axis.

Parameters:

step_data (np.ndarray) – Per-row step identifier (e.g. the "Step count" column).

Returns:

Index of the most recent step-start row for each row, same length as step_data.

Return type:

np.ndarray

ionworks.validators.running_step_reset_integral(signed, time, step_data=None, *, boundary_idx=None)[source]#

Cumulative trapezoidal integral of signed that resets at each step.

Mirrors the reset semantics of the platform’s cumulative capacity and energy columns so the returned series can be compared row-by-row. Divides by 3600 so the result is in A.h when signed is in A, or in W.h when signed is in W.

Parameters:
  • signed (np.ndarray) – Per-row signed values (e.g. max(I, 0) to integrate the discharge half-wave only).

  • time (np.ndarray) – Time values [s]; same length as signed.

  • step_data (np.ndarray, optional) – Per-row step identifier (e.g. "Step count" column). The integral is reset to 0 at every change in this value. When None (and boundary_idx is also None), the integral is not reset.

  • boundary_idx (np.ndarray, optional) – Output of step_boundary_idx() for the same step axis. Provide this to amortise the boundary computation across multiple integrals that share the same step_data.

Returns:

Running per-step integral, same length as signed, reset to 0 at the first row of each step.

Return type:

np.ndarray

ionworks.validators.worst_row_relative_error(reported, integrated)[source]#

Index, magnitude, and scale of the largest row-wise relative deviation.

The relative error is the largest absolute deviation between the two series divided by the larger of their max-magnitudes (the scale). NaN rows are ignored. Returns (0, inf, 0.0) when there is nothing to compare (both series flat at zero, or all-NaN).

Parameters:
  • reported (np.ndarray) – Reported cumulative series.

  • integrated (np.ndarray) – Integrated series to compare against, same length as reported.

Returns:

(worst_row_index, relative_error, max_scale).

Return type:

tuple[int, float, float]

ionworks.validators.validate_capacity_energy_from_current_power(df, step_col='Step count', time_col='Time [s]', current_col='Current [A]', voltage_col='Voltage [V]', power_col='Power [W]', discharge_capacity_col='Discharge capacity [A.h]', charge_capacity_col='Charge capacity [A.h]', discharge_energy_col='Discharge energy [W.h]', charge_energy_col='Charge energy [W.h]', tolerance=0.1)[source]#

Validate cumulative capacity/energy columns match integrals of current/power.

Builds a running trapezoidal integral over the whole time series and compares it row-by-row against the reported cumulative columns. The reported columns reset to 0 at the start of each step, so the integral is reset the same way and the comparison is made on the running within-step accumulator. With the platform sign convention (positive current = discharge):

  • Discharge capacity [A.h] ≈ ∫ max(I, 0) dt / 3600

  • Charge capacity [A.h] ≈ ∫ max(-I, 0) dt / 3600

  • Discharge energy [W.h] ≈ ∫ max(P, 0) dt / 3600

  • Charge energy [W.h] ≈ ∫ max(-P, 0) dt / 3600

The reported series and the integrated series are compared at every row. The relative error is the maximum absolute deviation across the whole series divided by the maximum cumulative value reached on either side, so a local discrepancy that later cancels out still triggers the check. A single issue is raised per cumulative column whose error exceeds tolerance. Columns missing from df are silently skipped.

Parameters:
  • df (DataFrame) – Time series data (pandas or polars).

  • step_col (str, optional) – Name of the step identifier column.

  • time_col (str, optional) – Name of the time column [s].

  • current_col (str, optional) – Name of the signed current column [A].

  • voltage_col (str, optional) – Name of the terminal voltage column [V]. Used to derive power as V * I when power_col is not present in df.

  • power_col (str, optional) – Name of the signed power column [W]. When absent, power is computed from voltage_col * current_col.

  • discharge_capacity_col (str, optional) – Names of the cumulative capacity columns [A.h].

  • charge_capacity_col (str, optional) – Names of the cumulative capacity columns [A.h].

  • discharge_energy_col (str, optional) – Names of the cumulative energy columns [W.h].

  • charge_energy_col (str, optional) – Names of the cumulative energy columns [W.h].

  • tolerance (float, optional) – Maximum allowed relative error between the integrated and reported series, evaluated as max|reported - integrated| / max(reported, integrated) across all rows. Defaults to 0.10 (10 %).

Returns:

One issue per cumulative column whose running series deviates from the integrated series by more than tolerance at any row. Empty when all present columns agree within tolerance everywhere.

Return type:

list[ValidationIssue]

ionworks.validators.validate_measurement_timing(start_time, end_time, df=None, *, duration_tolerance_frac=0.05, duration_tolerance_min_s=60.0, now=None)[source]#

Sanity-check a measurement’s wall-clock start_time / end_time.

All findings are warning severity — timing metadata is often approximate (clock skew, timezone sloppiness, planned runs), so these surface a nudge without blocking an upload. Checks:

  • start_time / end_time should not be in the future.

  • end_time should not precede start_time (the server enforces this as a hard error; this is an early client-side echo).

  • When df is given, the wall-clock duration end_time - start_time should roughly match the elapsed span in the Time [s] column (which is relative, starting at 0). A large mismatch suggests a mislabelled timestamp or a paused/segmented test.

Parameters:
  • start_time (datetime | str | None) – The measurement’s timestamps (ISO 8601 strings or datetime). A None skips the checks that need it (e.g. a still-running test has no end_time).

  • end_time (datetime | str | None) – The measurement’s timestamps (ISO 8601 strings or datetime). A None skips the checks that need it (e.g. a still-running test has no end_time).

  • df (DataFrame | None, optional) – Time series with a Time [s] column, used for the duration check.

  • duration_tolerance_frac (float, optional) – Allowed relative difference between wall-clock and data duration before warning. Defaults to 0.05 (5 %).

  • duration_tolerance_min_s (float, optional) – Absolute floor for the tolerance in seconds, so short tests aren’t flagged by rounding. Defaults to 60 s.

  • now (datetime | None, optional) – Reference “now” for the future checks; defaults to the current UTC time. Injectable for deterministic tests.

Returns:

Warning-severity issues (possibly empty).

Return type:

list[ValidationIssue]

ionworks.validators.validate_measurement_data(df, strict=False, data_type=None, steps_df=None, rated_capacity=None, voltage_window=None, skip_checks=None, start_time=None, end_time=None)[source]#

Validate measurement time series data before upload.

For standard cycler data (data_type=None), always runs:

  1. Positive current should correspond to discharge (voltage decreases)

  2. Time starts at 0

  3. Time is monotonically non-decreasing

  4. ‘Step count’ column exists, starts at 0, and increases by 1

  5. Cumulative values (capacity, energy) reset at each step start and only increase within steps

The remaining checks are strict-mode only (strict=True):

  1. Each step has at least 2 data points

  2. Cycle number does not change within a step

  3. No time gap between consecutive rows exceeds 5 hours

  4. When voltage_window is provided, voltage is continuous between consecutive rows (no systematic chronological reordering)

  5. When steps_df and rated_capacity are provided, two consecutive steps delivering more than 2× rated capacity do not share the same direction

  6. When steps_df and rated_capacity are provided, no single step exceeds 500 % of the rated capacity (soft warning)

  7. Reported cumulative capacity/energy columns agree with the trapezoidal integral of current/power within 10 %

  8. When steps_df is provided, each clearly-signed step accumulates its capacity/energy in the column matching its current direction — catches charge/discharge column pairs whose labels are swapped

For OCP data (data_type="ocp"), only validates:

  1. ‘Voltage [V]’ column exists

  2. ‘Step count’ column exists and is sequential

For EIS data (data_type="eis"), always validates:

  1. The ‘Frequency [Hz]’, ‘Z_Re [Ohm]’, and ‘Z_Im [Ohm]’ columns exist

and, in strict mode only (both skippable):

  1. ‘Z_Im [Ohm]’ sign is not reversed (capacitive band should be negative in the raw Im(Z) convention)

  2. When rated_capacity is provided, the ohmic resistance R0 = min(Z_Re) lies within a wide (100×) plausibility band around a capacity-derived estimate — catches order-of-magnitude unit errors

Parameters:
  • df (DataFrame) – Time series data to validate (pandas or polars DataFrame).

  • strict (bool) – If False (default), run only the always-on checks above. If True, additionally run: minimum 2 points per step, cycle number constant within step, time-gap check, voltage-continuity check (when voltage_window is provided), and the step-capacity checks (when steps_df and rated_capacity are provided).

  • data_type (str | None) – The type of data being validated. Use "ocp" for open-circuit potential data, which relaxes validation to skip current, time, capacity, and energy checks. Use "eis" for impedance spectra, which validates the EIS columns and (strict-only) the Z_Im sign and impedance magnitude instead of the cycler checks. Default is None (standard cycler data).

  • steps_df (DataFrame | None) – Optional step summary dataframe. Used only in strict mode for the charge/discharge column-direction check and — together with rated_capacity — the consecutive same-direction and per-step capacity checks.

  • rated_capacity (float | None) – Rated (nominal) cell capacity in A.h. Used only in strict mode: with steps_df it enables the consecutive-full-step and per-step capacity soft warnings; with data_type="eis" it enables the EIS impedance-magnitude sanity check.

  • voltage_window (tuple[float, float] | None) – Rated (V_min, V_max) voltage window of the cell. Used only in strict mode to enable the voltage-continuity check.

  • start_time (datetime | str | None) – The measurement’s wall-clock timestamps. When either is provided, a set of warning-severity timing sanity checks runs (never blocks the upload): a future start_time/end_time, an end_time before start_time (which the server rejects as a hard error), and — when both are present with the time series — a wall-clock duration that disagrees with the Time [s] span. See validate_measurement_timing().

  • end_time (datetime | str | None) – The measurement’s wall-clock timestamps. When either is provided, a set of warning-severity timing sanity checks runs (never blocks the upload): a future start_time/end_time, an end_time before start_time (which the server rejects as a hard error), and — when both are present with the time series — a wall-clock duration that disagrees with the Time [s] span. See validate_measurement_timing().

  • skip_checks (Iterable[str] | None) – Names of strict-mode checks to skip while keeping strict=True for everything else. Use this to relax a single known-problematic check (least-privilege) instead of disabling strict mode entirely. Recognized names are listed in STRICT_CHECK_NAMES: "minimum_points_per_step", "cycle_constant_within_step", "time_gaps", "voltage_continuity", "consecutive_same_direction_full_steps", "step_capacity_within_rated", "charge_discharge_column_direction", "capacity_energy_from_current_power", "eis_zim_sign", "eis_impedance_magnitude". Unknown names raise ValueError.

Raises:

MeasurementValidationError – If any validation checks fail. The exception contains a list of all errors found.

Return type:

None

ionworks.validators.df_to_dict_validator(v)[source]#

Convert DataFrame to dict with orient=’list’ for serialization.

Parameters:

v (Any)

Return type:

Any

ionworks.validators.dict_to_df_validator(v, return_type=None)[source]#

Convert dict to DataFrame for data processing.

Parameters:
  • v (Any) – Value to convert. If dict, converts to DataFrame.

  • return_type (str | None) – Type of DataFrame to return: “polars” or “pandas”. If None, uses the global setting from set_dataframe_backend().

Returns:

DataFrame if input was dict, otherwise unchanged.

Return type:

Any

ionworks.validators.parameter_validator(v)[source]#

Convert pybamm.Symbol values to JSON-serializable form.

Parameters:

v (Any)

Return type:

Any

ionworks.validators.float_sanitizer(v)[source]#

Sanitize float values to JSON-compatible forms.

Converts inf, -inf, and NaN to None since these are not JSON-compliant.

Parameters:

v (Any)

Return type:

Any

ionworks.validators.bounds_tuple_validator(v)[source]#

Convert bounds 2-tuple to list for JSON serialization.

Parameters:

v (Any) – Value to validate. If it’s a tuple with 2 elements, converts to list.

Returns:

List if input was a 2-tuple, otherwise unchanged.

Return type:

Any

ionworks.validators.file_scheme_validator(v)[source]#

Convert file:// and folder:// scheme paths to serialized dicts.

Handles file: prefixed paths (loads CSV as dict) and folder: prefixed paths (loads time_series and steps as dict). For folder:, parquet files are preferred over CSV when both are present. All other values are returned unchanged.

The referenced contents are read from the local filesystem and inlined into the submitted config, so the time series is subject to the same 1,000-row inline limit as a bare DataFrame; larger datasets must be uploaded as a measurement and referenced with db:<measurement_id>.

Raises:
Parameters:

v (Any)

Return type:

Any

ionworks.validators.pybamm_model_validator(v)[source]#

Convert pybamm.BaseModel instances to JSON-serializable config dicts.

Parameters:

v (Any)

Return type:

Any

ionworks.validators.run_validators_outbound(v)[source]#

Recursively apply outbound validators to values and nested containers.

Parameters:

v (Any)

Return type:

Any

ionworks.validators.run_validators_inbound(v)[source]#

Recursively apply inbound validators to values and nested containers.

Parameters:

v (Any)

Return type:

Any