Validators#
Outbound and inbound validation pipelines for data transformation.
Reusable validator functions and composable pipelines for value normalization.
Provides functions for composable inbound/outbound value normalization (e.g., converting between pandas DataFrames and dictionaries).
- ionworks.validators.set_dataframe_backend(backend)[source]#
Set the default DataFrame backend for data fetching.
This overrides the IONWORKS_DATAFRAME_BACKEND environment variable.
- Parameters:
backend (str) – DataFrame backend to use: “polars” or “pandas”.
- Raises:
ValueError – If backend is not “polars” or “pandas”.
- Return type:
None
- ionworks.validators.get_dataframe_backend()[source]#
Get the current DataFrame backend setting.
- Returns:
Current backend: “polars” or “pandas”.
- Return type:
- ionworks.validators.CAPACITY_ENERGY_INTEGRAL_TOLERANCE: float = 0.1#
Default relative-error tolerance for capacity/energy integral checks. Used by
validate_capacity_energy_from_current_power()and byionworksdata.transform.fix_swapped_charge_discharge_columnsso both agree on the same definition of “within tolerance”.
- class ionworks.validators.IssueCode(*values)[source]#
Bases:
StrEnumStable identifiers for measurement validation findings.
Use these — not substrings of the human-readable
message— when branching on issue identity. Members arestr-compatible.- CURRENT_SIGN_REVERSED = 'current_sign_reversed'#
- CURRENT_SIGN_INDETERMINATE = 'current_sign_indeterminate'#
- CURRENT_SIGN_UNSIGNED = 'current_sign_unsigned'#
- CUMULATIVE_VALUE_NOT_RESET = 'cumulative_value_not_reset'#
- CUMULATIVE_VALUE_DECREASED = 'cumulative_value_decreased'#
- STEP_TOO_FEW_POINTS = 'step_too_few_points'#
- TIME_DOES_NOT_START_AT_ZERO = 'time_does_not_start_at_zero'#
- TIME_NOT_MONOTONIC = 'time_not_monotonic'#
- STEP_COUNT_MISSING = 'step_count_missing'#
- STEP_COUNT_DOES_NOT_START_AT_ZERO = 'step_count_does_not_start_at_zero'#
- STEP_COUNT_NON_SEQUENTIAL = 'step_count_non_sequential'#
- CYCLE_CHANGES_WITHIN_STEP = 'cycle_changes_within_step'#
- OCP_VOLTAGE_COLUMN_MISSING = 'ocp_voltage_column_missing'#
- OCP_X_AXIS_COLUMN_MISSING = 'ocp_x_axis_column_missing'#
- TIME_SERIES_ROW_COUNT_EXCEEDED = 'time_series_row_count_exceeded'#
- TIME_GAP_TOO_LARGE = 'time_gap_too_large'#
- VOLTAGE_CONTINUITY = 'voltage_continuity'#
- CONSECUTIVE_SAME_DIRECTION_FULL_STEPS = 'consecutive_same_direction_full_steps'#
- STEP_CAPACITY_EXCEEDS_RATED = 'step_capacity_exceeds_rated'#
- CHARGE_DISCHARGE_CAPACITY_COLUMNS_SWAPPED = 'charge_discharge_capacity_columns_swapped'#
- CHARGE_DISCHARGE_ENERGY_COLUMNS_SWAPPED = 'charge_discharge_energy_columns_swapped'#
- DISCHARGE_CAPACITY_INTEGRAL_MISMATCH = 'discharge_capacity_integral_mismatch'#
- CHARGE_CAPACITY_INTEGRAL_MISMATCH = 'charge_capacity_integral_mismatch'#
- DISCHARGE_ENERGY_INTEGRAL_MISMATCH = 'discharge_energy_integral_mismatch'#
- CHARGE_ENERGY_INTEGRAL_MISMATCH = 'charge_energy_integral_mismatch'#
- EIS_COLUMNS_MISSING = 'eis_columns_missing'#
- EIS_ZIM_SIGN_REVERSED = 'eis_zim_sign_reversed'#
- EIS_IMPEDANCE_MAGNITUDE_IMPLAUSIBLE = 'eis_impedance_magnitude_implausible'#
- START_TIME_IN_FUTURE = 'start_time_in_future'#
- END_TIME_IN_FUTURE = 'end_time_in_future'#
- END_TIME_BEFORE_START_TIME = 'end_time_before_start_time'#
- DURATION_MISMATCH = 'duration_mismatch'#
- class ionworks.validators.ValidationIssue(code, message, severity='error', payload=<factory>)[source]#
Bases:
objectStructured measurement validation finding.
Returned by the
validate_*functions and carried byMeasurementValidationError. Branch oncoderather than parsingmessage; messages are for display only and may be reworded between releases.- Parameters:
code (IssueCode) – Stable identifier of the finding.
message (str) – Human-readable description of the failure, including any fix hint.
severity ({"error", "warning"}, optional) –
"error"(the default) causesvalidate_measurement_data()to raise."warning"is emitted viawarnings.warn()and is not attached to the resultingMeasurementValidationError— callers who need programmatic access to warning-severity issues must invoke the underlyingvalidate_*producer directly rather than relying one.has_code(...).payload (dict, optional) – Structured details: step indices, column names, observed values, thresholds. Keys depend on
code.
- exception ionworks.validators.MeasurementValidationError(message, errors=None)[source]#
Bases:
IonworksErrorException raised when measurement data validation fails.
errorsis a list of structuredValidationIssuerecords. Branch onissue.coderather than parsingissue.message.- Parameters:
message (str)
errors (list[ValidationIssue])
- Return type:
None
- __init__(message, errors=None)[source]#
Initialize the IonworksError.
- Parameters:
message (str | dict[str, Any]) – Error message string or dict containing error details. Supports both the legacy
{"detail": ...}format and the new standardized{"error_code": ..., "message": ..., "detail": ...}format.status_code (int | None) – Optional HTTP status code.
errors (list[ValidationIssue] | None)
- Return type:
None
- errors: list[ValidationIssue]#
- ionworks.validators.positive_current_is_charge(t, current, voltage)[source]#
Determine whether positive current corresponds to charging.
Fits an OCV-R equivalent-circuit model
V = OCV(SOC) - I * R0under two sign-convention hypotheses (positive = discharge vs. positive = charge). OCV is linear in SOC, constrained to a non-negative slope (monotonically increasing); R0 is a single positive scalar. When the sign convention is wrong the SOC axis is inverted, forcing the OCV slope toward zero. The convention producing the larger OCV slope is selected.- Parameters:
t (np.ndarray) – Time values [s].
current (np.ndarray) – Current values [A].
voltage (np.ndarray) – Voltage values [V].
- Returns:
is_charge (bool) –
Trueif positive current is charging,Falseif discharging. ReturnsFalsewhen there is insufficient data.p_value (float) – Confidence metric in [0, 1]. Lower values indicate higher confidence. Computed as the sum of two ratios clipped to [0, 1]: the ratio of the smaller to the larger OCV slope (
delta_ratio) and the ratio of the winner’s SSE to the loser’s SSE (sse_ratio). Returns 1.0 when the result is ambiguous or data is insufficient.
- Return type:
- ionworks.validators.validate_positive_current_is_discharge(df, current_col='Current [A]', voltage_col='Voltage [V]', time_col='Time [s]', step_col=None, rest_tol=0.001, relative_rest_tol_frac=0.02, relative_rest_tol_percentile=95.0)[source]#
Validate that positive current corresponds to discharge.
Discharge should cause voltage to decrease. This function analyzes the relationship between current direction and voltage change to verify the sign convention is correct.
Fits an OCV-R ECM per step, then uses a confidence vote across steps weighted by the trapezoidal integral of
|I(t)| dtover each step so that steps actually moving charge dominate the decision and long near-zero-current voltage holds contribute negligible weight.- Parameters:
df (DataFrame) – Time series data with current and voltage columns (pandas or polars).
current_col (str) – Name of the current column.
voltage_col (str) – Name of the voltage column.
time_col (str) – Name of the time column.
step_col (str, optional) – Name of the step column. If provided, analyzes per-step. Otherwise, infers steps from current sign changes.
rest_tol (float) – Tolerance for considering current as zero (rest).
relative_rest_tol_frac (float) – Fraction of a robust current scale used to set a relative rest threshold. The robust scale is the percentile of
abs(current)given byrelative_rest_tol_percentile.relative_rest_tol_percentile (float) – Percentile in [0, 100] used to estimate the robust current scale for the relative threshold. The effective rest threshold is
min(rest_tol, relative_rest_tol)so the stricter (smaller) threshold takes precedence.
- Returns:
List of validation issues. Empty if validation passes.
- Return type:
- ionworks.validators.validate_cumulative_values_reset_per_step(df, step_col='Step count', cumulative_cols=None, tolerance=1e-06)[source]#
Validate cumulative values reset to ~0 at each step and only increase.
- Parameters:
df (DataFrame) – Time series data (pandas or polars).
step_col (str) – Name of the column containing step numbers.
cumulative_cols (list[str], optional) – List of cumulative column names to validate. If None, checks for common capacity and energy columns.
tolerance (float) – Tolerance for considering a value as “zero” at step start.
- Returns:
List of validation issues. Empty if validation passes.
- Return type:
- ionworks.validators.validate_minimum_points_per_step(df, step_col='Step count', min_points=2)[source]#
Validate that each step has at least a minimum number of data points.
- Parameters:
- Returns:
List of validation issues. Empty if validation passes.
- Return type:
- ionworks.validators.validate_time_starts_at_zero(df, tolerance=1e-06)[source]#
Validate that ‘Time [s]’ starts at 0.
- Parameters:
df (DataFrame) – Time series data (pandas or polars).
tolerance (float) – Tolerance for considering the start value as zero.
- Returns:
List of validation issues. Empty if validation passes.
- Return type:
- ionworks.validators.validate_time_monotonic(df, time_col='Time [s]', tolerance=1e-12)[source]#
Validate that the time column is monotonically non-decreasing.
- Parameters:
- Returns:
List of validation issues. Empty if validation passes.
- Return type:
- ionworks.validators.validate_step_count_sequential(df)[source]#
Validate that ‘Step count’ exists, starts at 0, and increases by 1.
- Parameters:
df (DataFrame) – Time series data (pandas or polars).
- Returns:
List of validation issues. Empty if validation passes.
- Return type:
- ionworks.validators.validate_cycle_constant_within_step(df, step_col='Step count', cycle_col=None)[source]#
Validate that cycle number does not change within a step.
- Parameters:
- Returns:
List of validation issues. Empty if validation passes.
- Return type:
- ionworks.validators.validate_ocp_columns(df)[source]#
Validate that OCP data has required columns.
Checks that the DataFrame contains: 1. A ‘Voltage [V]’ column 2. At least one x-axis column: ‘Capacity [A.h]’, ‘Stoichiometry’, or ‘SOC’
- Parameters:
df (DataFrame) – Time series data (pandas or polars).
- Returns:
List of validation issues. Empty if validation passes.
- Return type:
- ionworks.validators.EIS_FREQ_COL = 'Frequency [Hz]'#
Default EIS column names (BioLogic/Gamry reader convention).
- ionworks.validators.EIS_RESISTANCE_CAPACITY_PRODUCT: float = 0.05#
Roughly cell-size-invariant ohmic resistance-capacity product [Ohm.A.h], used to estimate a plausible ohmic-resistance scale from rated capacity alone (
R0 ~ product / capacity, since larger cells have proportionally more electrode area in parallel). For Li-ion this clusters around 0.05-0.09 Ohm.A.h (power cells lower, small high-energy cells higher; ~1 order of magnitude spread), so 0.05 is a central-but-conservative anchor. It is only an order-of-magnitude reference for the deliberately wide (default 100x) EIS magnitude sanity band, so the exact value is not load-bearing.
- ionworks.validators.validate_eis_columns(df, freq_col='Frequency [Hz]', zre_col='Z_Re [Ohm]', zim_col='Z_Im [Ohm]')[source]#
Validate that EIS data has the required frequency and impedance columns.
- Parameters:
df (DataFrame) – EIS spectrum data (pandas or polars).
freq_col (str, optional) – Name of the frequency column. Defaults to
"Frequency [Hz]".zre_col (str, optional) – Name of the real-impedance column. Defaults to
"Z_Re [Ohm]".zim_col (str, optional) – Name of the imaginary-impedance column. Defaults to
"Z_Im [Ohm]".
- Returns:
A single
EIS_COLUMNS_MISSINGissue listing the missing and available columns, or an empty list when all required columns are present.- Return type:
- ionworks.validators.validate_eis_zim_sign(df, freq_col='Frequency [Hz]', zre_col='Z_Re [Ohm]', zim_col='Z_Im [Ohm]', reversed_fraction_threshold=0.75)[source]#
Validate that the EIS
Z_Imsign convention is not reversed.The database stores the raw imaginary part
Z_Im = Im(Z), which is negative across the capacitive arc; only a possible high-frequency inductive loop is positive. This check sorts by frequency, trims the inductive tail (points above theR0 = min(Z_Re)intercept), and flags the spectrum when the capacitive band is predominantly positive — i.e. the whole spectrum has been stored with a flipped sign. Points are weighted by|Z_Im|so the near-zero points around the real-axis crossing do not dominate the vote.- Parameters:
df (DataFrame) – EIS spectrum data (pandas or polars).
freq_col (str, optional) – Name of the frequency column. Defaults to
"Frequency [Hz]".zre_col (str, optional) – Name of the real-impedance column. Defaults to
"Z_Re [Ohm]".zim_col (str, optional) – Name of the imaginary-impedance column. Defaults to
"Z_Im [Ohm]".reversed_fraction_threshold (float, optional) – Minimum
|Z_Im|-weighted fraction of the capacitive band that must be positive to flag a reversed sign. Defaults to0.75(75 %).
- Returns:
A single
EIS_ZIM_SIGN_REVERSEDissue when the sign appears flipped, or an empty list otherwise (including when columns are missing or there are too few points).- Return type:
- ionworks.validators.validate_eis_impedance_magnitude(df, rated_capacity, factor=100.0, r_times_q=0.05, zre_col='Z_Re [Ohm]')[source]#
Check that the EIS ohmic resistance is within a wide plausibility band.
Estimates a plausible ohmic-resistance scale from the rated capacity alone via the roughly cell-size-invariant product
R0 ~ r_times_q / rated_capacity, then flags the spectrum only when the measured interceptR0 = min(Z_Re)is more thanfactortimes away from it in either direction. The band is deliberately wide (default100x) so it fires only on order-of-magnitude errors — e.g. an impedance stored in the wrong unit — and never on a plausible spectrum. The exactr_times_qvalue is unimportant at this tolerance.- Parameters:
df (DataFrame) – EIS spectrum data (pandas or polars).
rated_capacity (float | None) – Rated (nominal) cell capacity in A.h. The check is skipped when this is
Noneor non-positive.factor (float, optional) – Half-width of the sanity band as a multiplicative factor around the expected ohmic resistance. Defaults to
100.0.r_times_q (float, optional) – Order-of-magnitude resistance-capacity product [Ohm.A.h] used to estimate the expected ohmic resistance. Defaults to
EIS_RESISTANCE_CAPACITY_PRODUCT.zre_col (str, optional) – Name of the real-impedance column. Defaults to
"Z_Re [Ohm]".
- Returns:
A single
EIS_IMPEDANCE_MAGNITUDE_IMPLAUSIBLEissue whenR0is outside the band, or an empty list otherwise (including when the rated capacity is unavailable, the column is missing, orR0is non-positive).- Return type:
- ionworks.validators.validate_time_series_row_count(df, max_rows=1000)[source]#
Validate that the time series does not exceed the maximum row count.
Datasets larger than
max_rowsshould be uploaded via the standard upload flow and then referenced with"db:<measurement_id>"oriwdata.DataLoader.from_db(MEASUREMENT_ID)in pipeline configurations.- Parameters:
df (DataFrame) – Time series data (pandas or polars).
max_rows (int) – Maximum allowed number of rows.
- Returns:
List of validation issues. Empty if validation passes.
- Return type:
- ionworks.validators.validate_time_gaps(df, time_col='Time [s]', max_gap_seconds=18000)[source]#
Validate that there are no large gaps between consecutive time samples.
A gap longer than
max_gap_secondsbetween two consecutive rows is almost always a sign that a chunk of cycling was dropped — for example, a rest period was recorded as elapsed time in the file but the intermediate rows were stripped, leaving the capacity integral to count current across the unrecorded interval.- Parameters:
- Returns:
List of validation issues. Empty if validation passes.
- Return type:
- ionworks.validators.validate_voltage_continuity(df, voltage_window, voltage_col='Voltage [V]', jump_fraction=0.8, max_bad_fraction=0.05)[source]#
Validate that consecutive rows do not exhibit unphysical voltage jumps.
Computes the absolute voltage difference between every pair of consecutive rows and counts the fraction that exceed
jump_fractionof the rated voltage window(V_max - V_min). If more thanmax_bad_fractionof row pairs exceed that threshold, the data is flagged as likely being out of chronological order (for example, after a faultypartition_by("Cycle_raw")grouping that interleaves pulse and rest rows from different cycles).- Parameters:
df (DataFrame) – Time series data with a voltage column (pandas or polars).
voltage_window (tuple[float, float]) –
(V_min, V_max)rated voltage window of the cell, typically the lower/upper cutoff voltages. Only the spanV_max - V_minis used.voltage_col (str, optional) – Name of the voltage column. Defaults to
"Voltage [V]".jump_fraction (float, optional) – Fraction of the voltage window considered the maximum physically plausible single-row voltage change. Defaults to
0.80(80 %).max_bad_fraction (float, optional) – Maximum fraction of consecutive row pairs allowed to exceed the jump threshold before the check fails. Defaults to
0.05(5 %).
- Returns:
List of validation issues. Empty if validation passes.
- Return type:
- ionworks.validators.validate_consecutive_same_direction_full_steps(steps_df, rated_capacity, step_type_col='Step type', discharge_capacity_col='Discharge capacity [A.h]', charge_capacity_col='Charge capacity [A.h]', step_count_col='Step count', full_step_fraction=2.0)[source]#
Validate that no two consecutive full-capacity steps share the same direction.
Walks the step summary in order. For each constant-current step (identified by its
Step type) that delivers more thanfull_step_fractionof the rated capacity, tracks the direction. If two such steps appear consecutively with the same direction (ignoring rest / EIS / unknown steps, which do not reset the streak), the check fails. This catches measurements that bundle multiple independent experiments — e.g. a discharge-rate file concatenating several CC discharges into a single measurement.- Parameters:
steps_df (DataFrame) – Step summary dataframe (as produced by
ionworksdata.steps.summarize).rated_capacity (float | None) – Rated (nominal) cell capacity in A.h. Used together with
full_step_fractionto decide whether a step counts as a full charge / discharge. The check is skipped when this isNoneor non-positive.step_type_col (str, optional) – Name of the column containing the step type labels (
"Rest","Constant current discharge", etc.). Defaults to"Step type".discharge_capacity_col (str, optional) – Name of the per-step discharge capacity column.
charge_capacity_col (str, optional) – Name of the per-step charge capacity column.
step_count_col (str, optional) – Name of the per-step identifier column, reported in error messages.
full_step_fraction (float, optional) – Multiple of rated capacity above which a step is considered a full (and then some) charge / discharge. Defaults to
2.0— i.e. only steps delivering more than 2× rated capacity trigger the check, which tolerates one or two full cycles being captured in a single step while still catching runaway concatenations.
- Returns:
List of validation issues. Empty if validation passes.
- Return type:
- ionworks.validators.validate_step_capacity_within_rated(steps_df, rated_capacity, discharge_capacity_col='Discharge capacity [A.h]', charge_capacity_col='Charge capacity [A.h]', step_count_col='Step count', max_ratio=5.0)[source]#
Soft-check that no single step exceeds
max_ratio× rated capacity.A single step accumulating more capacity than several full charges or discharges of the cell usually indicates wrong step boundaries or that the capacity integral was inflated across unrecorded time gaps. Returns a list of warnings; callers may emit them via
warnings.warn()instead of raising.- Parameters:
steps_df (DataFrame) – Step summary dataframe (as produced by
ionworksdata.steps.summarize).rated_capacity (float | None) – Rated (nominal) cell capacity in A.h. The check is skipped when this is
Noneor non-positive.discharge_capacity_col (str, optional) – Name of the per-step discharge capacity column.
charge_capacity_col (str, optional) – Name of the per-step charge capacity column.
step_count_col (str, optional) – Name of the per-step identifier column, reported in warning messages.
max_ratio (float, optional) – Maximum allowed ratio of per-step capacity to rated capacity. Defaults to
5.0(500 %).
- Returns:
List of validation issues. Empty if no step exceeds the threshold.
- Return type:
- ionworks.validators.validate_charge_discharge_column_direction(steps_df, mean_current_col='Mean current [A]', min_current_col='Min current [A]', max_current_col='Max current [A]', std_current_col='Std current [A]', duration_col='Duration [s]', step_count_col='Step count', discharge_capacity_col='Discharge capacity [A.h]', charge_capacity_col='Charge capacity [A.h]', discharge_energy_col='Discharge energy [W.h]', charge_energy_col='Charge energy [W.h]', rest_tol=0.001, min_duration_s=60.0, max_std_to_mean_ratio=0.5, dominance_ratio=4.0, swapped_fraction_threshold=0.75)[source]#
Detect swapped charge/discharge capacity (and energy) columns in a step summary.
With the platform sign convention (positive current = discharge), a step that only discharges must accumulate its A.h in the discharge column, and a step that only charges must accumulate in the charge column. A cycler export with the two column labels inverted is symmetric under every cumulative-reset and magnitude check, so this is the only signal that catches it: for each step whose current direction is unambiguous, compare the sign of the mean current against which column of the pair actually accumulated, and flag the pair when the clear majority of such steps put their capacity in the opposite-direction column.
A step votes only when the evidence is unambiguous on every axis:
its mean current is clearly non-zero (
|mean| >= rest_tol),it lasted at least
min_duration_s(when a duration column is available),its current did not change sign — the min/max currents stay on the same side of zero (within
rest_tol), or the standard deviation is small relative to|mean|; either piece of evidence qualifies, so a single opposite-sign boundary sample inherited from the preceding step does not silence an otherwise one-directional step. Bidirectional (drive-cycle) steps legitimately accumulate in both columns and are excluded by this test, andone column of the pair clearly dominates the other (
dominance_ratio), so a step that accumulated comparably in both columns never votes.
Votes are weighted by the capacity (or energy) the step moved, so a handful of tiny glitch steps cannot outvote real cycling. The capacity pair and the energy pair are evaluated independently; energy direction is also judged from the mean current, since terminal voltage is positive and power therefore shares the current’s sign.
- Parameters:
steps_df (DataFrame) – Step summary dataframe (as produced by
ionworksdata.steps.identify), pandas or polars.mean_current_col (str, optional) – Name of the per-step mean current column [A]. The check is skipped entirely when this column is absent.
min_current_col (str, optional) – Names of the per-step min/max current columns [A], used to establish that a step’s current never changed sign.
max_current_col (str, optional) – Names of the per-step min/max current columns [A], used to establish that a step’s current never changed sign.
std_current_col (str, optional) – Name of the per-step current standard deviation column [A]. Alternative sign-unambiguity evidence: the step also qualifies when
std <= max_std_to_mean_ratio * |mean|, even when its min/max fail the same-sign test (e.g. polluted by a boundary sample).duration_col (str, optional) – Name of the per-step duration column [s]. Steps shorter than
min_duration_s(or with NaN duration) do not vote. No duration filter is applied when the column is absent.step_count_col (str, optional) – Name of the per-step identifier column, reported in the issue payload.
discharge_capacity_col (str, optional) – Names of the per-step capacity pair [A.h]. The pair is skipped when either column is absent.
charge_capacity_col (str, optional) – Names of the per-step capacity pair [A.h]. The pair is skipped when either column is absent.
discharge_energy_col (str, optional) – Names of the per-step energy pair [W.h]. The pair is skipped when either column is absent.
charge_energy_col (str, optional) – Names of the per-step energy pair [W.h]. The pair is skipped when either column is absent.
rest_tol (float, optional) – Current magnitude [A] below which a step counts as rest and does not vote. Also the tolerance for the min/max same-sign test, so a discharge step whose current briefly touches zero still qualifies. Defaults to
1e-3.min_duration_s (float, optional) – Minimum step duration [s] for a step to vote. Defaults to
60.max_std_to_mean_ratio (float, optional) – Maximum
std / |mean|for a step to qualify via the standard-deviation fallback. Defaults to0.5.dominance_ratio (float, optional) – Minimum ratio of the larger to the smaller column value for the step’s accumulation to count as clearly one-directional. Defaults to
4.0.swapped_fraction_threshold (float, optional) – Minimum weighted fraction of voting steps that must accumulate in the wrong column to flag the pair as swapped. Defaults to
0.75.
- Returns:
At most one issue per pair (
CHARGE_DISCHARGE_CAPACITY_COLUMNS_SWAPPEDand/orCHARGE_DISCHARGE_ENERGY_COLUMNS_SWAPPED). Empty when neither pair appears swapped or there is not enough unambiguous evidence.- Return type:
- ionworks.validators.step_boundary_idx(step_data)[source]#
Per-row index of the most recent step start.
Subtracting
series[step_boundary_idx(step_data)]from a running cumulativeseriesresets it to 0 at every change instep_data. Compute this once and pass it into multiplerunning_step_reset_integral()calls when they share the same step axis.- Parameters:
step_data (np.ndarray) – Per-row step identifier (e.g. the
"Step count"column).- Returns:
Index of the most recent step-start row for each row, same length as
step_data.- Return type:
np.ndarray
- ionworks.validators.running_step_reset_integral(signed, time, step_data=None, *, boundary_idx=None)[source]#
Cumulative trapezoidal integral of
signedthat resets at each step.Mirrors the reset semantics of the platform’s cumulative capacity and energy columns so the returned series can be compared row-by-row. Divides by 3600 so the result is in A.h when
signedis in A, or in W.h whensignedis in W.- Parameters:
signed (np.ndarray) – Per-row signed values (e.g.
max(I, 0)to integrate the discharge half-wave only).time (np.ndarray) – Time values [s]; same length as
signed.step_data (np.ndarray, optional) – Per-row step identifier (e.g.
"Step count"column). The integral is reset to 0 at every change in this value. WhenNone(andboundary_idxis also None), the integral is not reset.boundary_idx (np.ndarray, optional) – Output of
step_boundary_idx()for the same step axis. Provide this to amortise the boundary computation across multiple integrals that share the samestep_data.
- Returns:
Running per-step integral, same length as
signed, reset to 0 at the first row of each step.- Return type:
np.ndarray
- ionworks.validators.worst_row_relative_error(reported, integrated)[source]#
Index, magnitude, and scale of the largest row-wise relative deviation.
The relative error is the largest absolute deviation between the two series divided by the larger of their max-magnitudes (the scale). NaN rows are ignored. Returns
(0, inf, 0.0)when there is nothing to compare (both series flat at zero, or all-NaN).
- ionworks.validators.validate_capacity_energy_from_current_power(df, step_col='Step count', time_col='Time [s]', current_col='Current [A]', voltage_col='Voltage [V]', power_col='Power [W]', discharge_capacity_col='Discharge capacity [A.h]', charge_capacity_col='Charge capacity [A.h]', discharge_energy_col='Discharge energy [W.h]', charge_energy_col='Charge energy [W.h]', tolerance=0.1)[source]#
Validate cumulative capacity/energy columns match integrals of current/power.
Builds a running trapezoidal integral over the whole time series and compares it row-by-row against the reported cumulative columns. The reported columns reset to 0 at the start of each step, so the integral is reset the same way and the comparison is made on the running within-step accumulator. With the platform sign convention (positive current = discharge):
Discharge capacity [A.h] ≈ ∫ max(I, 0) dt / 3600
Charge capacity [A.h] ≈ ∫ max(-I, 0) dt / 3600
Discharge energy [W.h] ≈ ∫ max(P, 0) dt / 3600
Charge energy [W.h] ≈ ∫ max(-P, 0) dt / 3600
The reported series and the integrated series are compared at every row. The relative error is the maximum absolute deviation across the whole series divided by the maximum cumulative value reached on either side, so a local discrepancy that later cancels out still triggers the check. A single issue is raised per cumulative column whose error exceeds
tolerance. Columns missing fromdfare silently skipped.- Parameters:
df (DataFrame) – Time series data (pandas or polars).
step_col (str, optional) – Name of the step identifier column.
time_col (str, optional) – Name of the time column [s].
current_col (str, optional) – Name of the signed current column [A].
voltage_col (str, optional) – Name of the terminal voltage column [V]. Used to derive power as
V * Iwhenpower_colis not present indf.power_col (str, optional) – Name of the signed power column [W]. When absent, power is computed from
voltage_col * current_col.discharge_capacity_col (str, optional) – Names of the cumulative capacity columns [A.h].
charge_capacity_col (str, optional) – Names of the cumulative capacity columns [A.h].
discharge_energy_col (str, optional) – Names of the cumulative energy columns [W.h].
charge_energy_col (str, optional) – Names of the cumulative energy columns [W.h].
tolerance (float, optional) – Maximum allowed relative error between the integrated and reported series, evaluated as
max|reported - integrated| / max(reported, integrated)across all rows. Defaults to0.10(10 %).
- Returns:
One issue per cumulative column whose running series deviates from the integrated series by more than
toleranceat any row. Empty when all present columns agree within tolerance everywhere.- Return type:
- ionworks.validators.validate_measurement_timing(start_time, end_time, df=None, *, duration_tolerance_frac=0.05, duration_tolerance_min_s=60.0, now=None)[source]#
Sanity-check a measurement’s wall-clock
start_time/end_time.All findings are warning severity — timing metadata is often approximate (clock skew, timezone sloppiness, planned runs), so these surface a nudge without blocking an upload. Checks:
start_time/end_timeshould not be in the future.end_timeshould not precedestart_time(the server enforces this as a hard error; this is an early client-side echo).When
dfis given, the wall-clock durationend_time - start_timeshould roughly match the elapsed span in theTime [s]column (which is relative, starting at 0). A large mismatch suggests a mislabelled timestamp or a paused/segmented test.
- Parameters:
start_time (datetime | str | None) – The measurement’s timestamps (ISO 8601 strings or
datetime). ANoneskips the checks that need it (e.g. a still-running test has noend_time).end_time (datetime | str | None) – The measurement’s timestamps (ISO 8601 strings or
datetime). ANoneskips the checks that need it (e.g. a still-running test has noend_time).df (DataFrame | None, optional) – Time series with a
Time [s]column, used for the duration check.duration_tolerance_frac (float, optional) – Allowed relative difference between wall-clock and data duration before warning. Defaults to 0.05 (5 %).
duration_tolerance_min_s (float, optional) – Absolute floor for the tolerance in seconds, so short tests aren’t flagged by rounding. Defaults to 60 s.
now (datetime | None, optional) – Reference “now” for the future checks; defaults to the current UTC time. Injectable for deterministic tests.
- Returns:
Warning-severity issues (possibly empty).
- Return type:
- ionworks.validators.validate_measurement_data(df, strict=False, data_type=None, steps_df=None, rated_capacity=None, voltage_window=None, skip_checks=None, start_time=None, end_time=None)[source]#
Validate measurement time series data before upload.
For standard cycler data (
data_type=None), always runs:Positive current should correspond to discharge (voltage decreases)
Time starts at 0
Time is monotonically non-decreasing
‘Step count’ column exists, starts at 0, and increases by 1
Cumulative values (capacity, energy) reset at each step start and only increase within steps
The remaining checks are strict-mode only (
strict=True):Each step has at least 2 data points
Cycle number does not change within a step
No time gap between consecutive rows exceeds 5 hours
When
voltage_windowis provided, voltage is continuous between consecutive rows (no systematic chronological reordering)When
steps_dfandrated_capacityare provided, two consecutive steps delivering more than 2× rated capacity do not share the same directionWhen
steps_dfandrated_capacityare provided, no single step exceeds 500 % of the rated capacity (soft warning)Reported cumulative capacity/energy columns agree with the trapezoidal integral of current/power within 10 %
When
steps_dfis provided, each clearly-signed step accumulates its capacity/energy in the column matching its current direction — catches charge/discharge column pairs whose labels are swapped
For OCP data (
data_type="ocp"), only validates:‘Voltage [V]’ column exists
‘Step count’ column exists and is sequential
For EIS data (
data_type="eis"), always validates:The ‘Frequency [Hz]’, ‘Z_Re [Ohm]’, and ‘Z_Im [Ohm]’ columns exist
and, in strict mode only (both skippable):
‘Z_Im [Ohm]’ sign is not reversed (capacitive band should be negative in the raw
Im(Z)convention)When
rated_capacityis provided, the ohmic resistanceR0 = min(Z_Re)lies within a wide (100×) plausibility band around a capacity-derived estimate — catches order-of-magnitude unit errors
- Parameters:
df (DataFrame) – Time series data to validate (pandas or polars DataFrame).
strict (bool) – If False (default), run only the always-on checks above. If True, additionally run: minimum 2 points per step, cycle number constant within step, time-gap check, voltage-continuity check (when
voltage_windowis provided), and the step-capacity checks (whensteps_dfandrated_capacityare provided).data_type (str | None) – The type of data being validated. Use
"ocp"for open-circuit potential data, which relaxes validation to skip current, time, capacity, and energy checks. Use"eis"for impedance spectra, which validates the EIS columns and (strict-only) theZ_Imsign and impedance magnitude instead of the cycler checks. Default isNone(standard cycler data).steps_df (DataFrame | None) – Optional step summary dataframe. Used only in strict mode for the charge/discharge column-direction check and — together with
rated_capacity— the consecutive same-direction and per-step capacity checks.rated_capacity (float | None) – Rated (nominal) cell capacity in A.h. Used only in strict mode: with
steps_dfit enables the consecutive-full-step and per-step capacity soft warnings; withdata_type="eis"it enables the EIS impedance-magnitude sanity check.voltage_window (tuple[float, float] | None) – Rated
(V_min, V_max)voltage window of the cell. Used only in strict mode to enable the voltage-continuity check.start_time (datetime | str | None) – The measurement’s wall-clock timestamps. When either is provided, a set of warning-severity timing sanity checks runs (never blocks the upload): a future
start_time/end_time, anend_timebeforestart_time(which the server rejects as a hard error), and — when both are present with the time series — a wall-clock duration that disagrees with theTime [s]span. Seevalidate_measurement_timing().end_time (datetime | str | None) – The measurement’s wall-clock timestamps. When either is provided, a set of warning-severity timing sanity checks runs (never blocks the upload): a future
start_time/end_time, anend_timebeforestart_time(which the server rejects as a hard error), and — when both are present with the time series — a wall-clock duration that disagrees with theTime [s]span. Seevalidate_measurement_timing().skip_checks (Iterable[str] | None) – Names of strict-mode checks to skip while keeping
strict=Truefor everything else. Use this to relax a single known-problematic check (least-privilege) instead of disabling strict mode entirely. Recognized names are listed inSTRICT_CHECK_NAMES:"minimum_points_per_step","cycle_constant_within_step","time_gaps","voltage_continuity","consecutive_same_direction_full_steps","step_capacity_within_rated","charge_discharge_column_direction","capacity_energy_from_current_power","eis_zim_sign","eis_impedance_magnitude". Unknown names raiseValueError.
- Raises:
MeasurementValidationError – If any validation checks fail. The exception contains a list of all errors found.
- Return type:
None
- ionworks.validators.df_to_dict_validator(v)[source]#
Convert DataFrame to dict with orient=’list’ for serialization.
- ionworks.validators.dict_to_df_validator(v, return_type=None)[source]#
Convert dict to DataFrame for data processing.
- Parameters:
v (Any) – Value to convert. If dict, converts to DataFrame.
return_type (str | None) – Type of DataFrame to return: “polars” or “pandas”. If None, uses the global setting from set_dataframe_backend().
- Returns:
DataFrame if input was dict, otherwise unchanged.
- Return type:
Any
- ionworks.validators.parameter_validator(v)[source]#
Convert pybamm.Symbol values to JSON-serializable form.
- ionworks.validators.float_sanitizer(v)[source]#
Sanitize float values to JSON-compatible forms.
Converts inf, -inf, and NaN to None since these are not JSON-compliant.
- ionworks.validators.bounds_tuple_validator(v)[source]#
Convert bounds 2-tuple to list for JSON serialization.
- Parameters:
v (Any) – Value to validate. If it’s a tuple with 2 elements, converts to list.
- Returns:
List if input was a 2-tuple, otherwise unchanged.
- Return type:
Any
- ionworks.validators.file_scheme_validator(v)[source]#
Convert file:// and folder:// scheme paths to serialized dicts.
Handles
file:prefixed paths (loads CSV as dict) andfolder:prefixed paths (loads time_series and steps as dict). Forfolder:, parquet files are preferred over CSV when both are present. All other values are returned unchanged.The referenced contents are read from the local filesystem and inlined into the submitted config, so the time series is subject to the same 1,000-row inline limit as a bare
DataFrame; larger datasets must be uploaded as a measurement and referenced withdb:<measurement_id>.- Raises:
FileNotFoundError – If the file or folder path doesn’t exist.
MeasurementValidationError – If the referenced time series exceeds the inline row limit.
- Parameters:
v (Any)
- Return type:
- ionworks.validators.pybamm_model_validator(v)[source]#
Convert pybamm.BaseModel instances to JSON-serializable config dicts.