Integration cookbook¶
Connect tslens to forecasting and classification models with the recipes below. Every snippet is exercised by the test suite or a runnable notebook.
The attribution contract¶
Every integration follows the same pattern:
attr = WinTSR(model).attribute(inputs, baselines=..., additional_forward_args=...)
inputs— tensors you want explained. They get perturbed, and you get one attribution per input.additional_forward_args— everything else the model needs. Passed through untouched, never attributed.
Attributions come back shaped (batch, n_output, seq_len, n_features).
| Recipe | Use it when |
|---|---|
| Your own model | The model accepts one time-series tensor. |
| TSlib models | The model accepts encoder and decoder tensors. |
| Classification | Model outputs are class scores. |
| Custom outputs | The model returns a dictionary or tuple. |
| Choosing a baseline | Zero is not an appropriate reference value. |
| WinIT reference data | Counterfactual values should come from real series. |
| GateMask training | You want a learned sparse mask and can afford an inner training loop. |
| Performance tuning | Attribution needs to be faster or sparser. |
| Troubleshooting | Shapes, outputs, or scores are unexpected. |
Your own model¶
If it maps (batch, seq_len, n_features) to predictions, there is nothing to configure.
import torch
from tslens import WinTSR
inputs = torch.randn(16, 96, 7)
attr = WinTSR(model).attribute(inputs, baselines=torch.zeros_like(inputs))
# (16, n_output, 96, 7)
Start with the quickstart notebook, then see the real-data case study for chronological splitting, training, held-out evaluation, and interpretation.
TSlib models (DLinear, iTransformer, TimesNet, Autoformer, ...)¶
TSlib models take four tensors. Split them: the two encoder inputs are attributed, the two decoder inputs are context. No wrapper class needed.
attr_enc, attr_mark = WinTSR(model).attribute(
inputs=(x_enc, x_mark_enc),
baselines=(torch.zeros_like(x_enc), torch.zeros_like(x_mark_enc)),
additional_forward_args=(x_dec, x_mark_dec),
threshold=0.5,
)
attr_enc is the one you usually want. Because the model returns
(batch, pred_len, c_out), n_output is pred_len * c_out — one saliency map per
predicted value.
Full walkthrough: TSlib models notebook.
Single-input foundation models (CALF, OFA/GPT4TS)¶
These consume only the series. Drop the tuples.
attr = WinTSR(model).attribute(
inputs=x_enc,
baselines=torch.zeros_like(x_enc),
)
Classification models¶
Pass the padding mask (and anything else) as context. n_output becomes the class count.
attr = WinTSR(model).attribute(
inputs=batch_x,
baselines=torch.zeros_like(batch_x),
additional_forward_args=(padding_mask,),
)
# (batch, n_classes, seq_len, n_features)
Full walkthrough: classification notebook.
Models that return a dict or a tuple¶
Attribution needs a tensor. Wrap the model in a callable that picks the right field —
WinTSR accepts any callable, not just nn.Module.
# model returns {"outputs_time": ..., "outputs_text": ...}
attr = WinTSR(lambda x: model(x)["outputs_time"]).attribute(inputs, baselines=...)
# model returns (predictions, attention_weights)
attr = WinTSR(lambda x: model(x)[0]).attribute(inputs, baselines=...)
Full walkthrough: custom outputs notebook.
Choosing a baseline¶
The baseline is what an occluded region is replaced with. Zeros are the default and are fine for standardized data.
from tslens import get_baseline
get_baseline(inputs, "zero") # zeros (default)
get_baseline(inputs, "random") # standard normal noise
get_baseline(inputs, "normal") # sampled from each feature's own mean/std
get_baseline(inputs, "mean") # each feature's mean, broadcast
Use "normal" or "mean" when zero is a meaningful value in your data and would itself
look like a signal.
Full walkthrough: baselines notebook.
WinIT reference data¶
WinIT samples counterfactual replacement values from a representative dataset rather
than accepting a fixed baseline. For forecasting, set task_name to anything other than
"classification" to use prediction difference:
import argparse
from tslens.attr import WinIT
args = argparse.Namespace(
seq_len=inputs.shape[1], task_name="forecast", pred_len=1, features="S", seed=0,
)
# WinIT expects forecasting outputs shaped (batch, pred_len, features).
winit_model = lambda x: model(x).unsqueeze(-1) # only if model(x) is (batch, pred_len)
attr = WinIT(winit_model, data=x_train, args=args).attribute(
inputs=inputs,
additional_forward_args=None,
attributions_fn=abs,
)
# (batch, n_output, seq_len, n_features)
Full walkthrough: WinIT notebook.
GateMask training¶
GateMask accepts one attributed input tensor and fits a mask network on each call. Its
output has the same (batch, seq_len, n_features) shape as the input, without WinTSR's
explicit n_output axis:
from pytorch_lightning import Trainer
from tslens.attr import GateMask
attr = GateMask(model).attribute(
inputs=inputs,
baselines=torch.zeros_like(inputs),
trainer=Trainer(
max_epochs=50, accelerator="cpu", enable_progress_bar=False,
enable_model_summary=False, logger=False,
),
batch_size=16,
)
Full walkthrough: GateMask notebook.
Explaining one forecast horizon¶
n_output indexes flattened predictions. For an output shaped
(batch, pred_len, channels), reshape that axis before selecting a horizon or channel:
attr = WinTSR(model).attribute(inputs, baselines=zeros)
attr_hc = attr.reshape(batch, pred_len, channels, seq_len, n_features)
attr_hc[:, 0] # first horizon, every predicted channel
attr_hc[:, :, ot_idx] # every horizon for predicted channel OT
attr.abs().mean(dim=1) # averaged over all outputs
The flattening is horizon-major: for native (horizon, channel) output, OT also appears
at attr[:, ot_idx::channels]. See the real-data case study
for this pattern on a 24-horizon, seven-channel forecast.
Performance tuning¶
threshold is the main dial. It is the quantile of time-relevance below which time steps
are skipped in stage two, so higher means fewer model calls and a sparser map.
WinTSR(model).attribute(inputs, baselines=zeros, threshold=0.0) # every time step
WinTSR(model).attribute(inputs, baselines=zeros, threshold=0.5) # skip the bottom half
WinTSR(model).attribute(inputs, baselines=zeros, threshold=0.9) # fastest, sparsest
Widening the temporal window also reduces the number of positions to evaluate:
WinTSR(model).attribute(inputs, baselines=zeros, sliding_window_shapes=(6, 1))
perturbations_per_eval > 1only works for single-output models. Multi-output models must leave it at1. This is an upstream limitation of tint's FeatureAblation — plaintint.attr.Occlusionfails the same way — and WinTSR raises a clearValueErrorexplaining it rather than letting the raw assertion through.
Comparing against other methods¶
WinTSR follows Captum's attribution interface, so comparisons are concise:
from captum.attr import IntegratedGradients
from tint.attr import FeatureAblation, Occlusion
from tslens.attr import WinTSR, WinIT, GateMask, TSR
Occlusion(model).attribute(inputs, sliding_window_shapes=(1, 1), baselines=zeros)
IntegratedGradients(model).attribute(inputs, baselines=zeros)
Troubleshooting¶
Start with these common failure modes. If the model has an unusual return type, see Models that return a dict or a tuple.
| Symptom | Cause |
|---|---|
AssertionError: All inputs must have the same time dimension |
Tensors in inputs disagree on shape[1]. Move the odd one to additional_forward_args. |
ValueError: perturbations_per_eval > 1 only works for single-output models |
Leave perturbations_per_eval at 1. |
| Attribution is entirely zero | threshold is too high; try 0.0. |
Shape is (batch * n_output, seq_len, n_features) |
You passed unflatten=False. |
| Attribution looks like noise | Check the model actually learned — an untrained model has nothing to explain. |