refactor(quickadapter): extract is_finite_number to Utils and reuse it (#122)
* refactor(quickadapter): extract is_finite_number to Utils and reuse it
- move the numeric/finite/non-bool scalar guard from a QuickAdapterV3
static method to a shared Utils.is_finite_number helper
- reuse it across the strategy label getters/setters and the shared
label_natr_multiplier validation guard
- leave the strict positive-int label_period_candles/label_horizon_candles
paths unchanged (distinct predicate)
* refactor(quickadapter): harden label-param fallbacks and per-row NATR guard
Address the #112 re-review findings (landed in this PR):
- coerce the per-row label_period_candles series with pandas to_numeric
before np.isfinite, matching the guarded scalar path (object/str dtype no
longer raises)
- re-validate the persisted _label_params fallback in the label getters via
is_finite_number before use, otherwise fall back to config
- generalize the is_trade_runmode comment (it gates both persisted-param
reuse and per-candle param setting)
- README: drop the inaccurate 'simulated' wording and the runmode-specific
framing on the label_period_candles/label_natr_multiplier rows
* docs(quickadapter): drop per-candle label-param behavior notes from tunables
The label_period_candles/label_natr_multiplier rows describe labeling
tunables; the per-candle strategy-NATR/exit consumption is an internal
behavior detail (already documented inline in the code), not needed to
configure the tunable. Revert both rows to their base description.
* style(quickadapter): wrap set_label_natr_multiplier guard per ruff format
The committed one-liner exceeded the 88-char line length; apply ruff
format wrapping (behavior unchanged).
* refactor(quickadapter): harden label-param validation and dedupe runmode gate
- add _is_finite_number guard rejecting bool and non-numeric before
np.isfinite (which raised on str/object) across the label getters/setters
- factor the duplicated runmode-in-TRADE_MODES predicate into the
is_trade_runmode cached_property
- document the per-candle NATR construction and drop the redundant fillna
after where() in the per-row NATR path
- note that backtest and hyperopt (not only backtest) retain the per-candle
label period in the README
fix(quickadapter): make HPO state causal in backtests (#111)
* fix(quickadapter): make label HPO causal in backtests
* fix(quickadapter): harden causal label HPO bounds, docs and harmonization
- guard empty prediction history and use NaT/order-safe max()/min() when
bounding label HPO OHLCV to the current FreqAI prediction time
- gate strategy label-param loading on freqtrade TRADE_MODES to mirror the
regressor self.live gate
- document the point-in-time study reset (supersedes explicit continuous=false)
and causal warm-start seeding
- make the optuna_hyperopt.enabled README entry terse and add a dedicated
causal label HPO note; unify terminology on current FreqAI prediction time
* docs(quickadapter): drop causal label HPO note from README
* fix(quickadapter): silence benign Optuna study deletion on fresh storage
- treat a missing study on delete as a debug no-op: non-live runs use a
fresh InMemoryStorage and the first live/dry-run optimization per pair
has no persisted study yet, so optuna.delete_study raises KeyError; keep
warning+traceback for genuine deletion failures
- drop redundant point-in-time frame copies: DataProvider.get_pair_dataframe
already returns a caller-owned frame and it is only read downstream
- lower the non-live 'Label HPO skipped' bounds logs to debug (expected
backtest warmup states, consistent with the throttle debug log)
- refine the point-in-time docstring (dk.full_df is the full feature frame)
and note that self.live is unset at __init__
* docs(quickadapter): tighten HPO comments and de-parenthesize README
- make the __init__ trade-mode, non-live point-in-time HPO, and delete_study
KeyError comments more precise and concise without dropping semantics
- reword the delete_study comment to point at the warning branch instead of
the inaccurate 'real failures raise otherwise'
- rephrase the optuna_hyperopt.continuous README entry without a parenthetical
precision, consistent with the surrounding prose
* chore(serena): migrate project config to the language_servers schema
Serena renamed the deprecated `languages` key to `language_servers` and
refreshed the accompanying comments; regenerate the tracked project config
to match.
chore(quickadapter): raise default fill_sigma_candles to 25.0
Update the canonical DEFAULTS_LABEL_WEIGHTING value, the config template,
and the tunables documentation in lockstep. The default is inert for the
default fill_method ("zero"); it only affects the per-pivot Gaussian
bandwidth (and its knn clip upper bound) when fill_method is "gaussian"
or "epsilon_gaussian".
refactor(quickadapter): disambiguate and centralize final exit stage constants
Rename the _FINAL_EXIT_STAGE tuple to _FINAL_EXIT_STAGE_PARAMS and add the
class-level constant _FINAL_EXIT_STAGE_INDEX (max(partial_exit_stages) + 1),
referenced from the configuration logging, the get_trade_exit_stage clamp,
and the plot config. This removes the triplicated stage-index expression and
distinguishes the final stage parameters from its index. Qualify the plot
config partial_exit_stages and stage-index accesses with QuickAdapterV3 to
match the surrounding constant access.
fix(quickadapter): bound take-profit stage to filled take-profit exits
Derive the take-profit stage from filled take-profit-tagged exit orders
clamped to the final full-exit stage, instead of nr_of_successful_exits
plus in-flight exit orders. This stops the stage index from overshooting
its maximum (e.g. take_profit_*_4 while the final stage is 3) when a
final exit order is in flight or a non-take-profit exit fills.
Guard custom_exit against re-issuing the final take-profit exit while an
order is open by returning None after the model-expiry and reversal
safety exits and before the take-profit stage computation, so safety
exits still fire while in-flight orders no longer trigger a duplicate
take-profit exit.
Jérôme Benoit [Mon, 22 Jun 2026 13:07:40 +0000 (15:07 +0200)]
style(quickadapter): emit label_horizon_candles per pair in HPO=on log block
Two post-implementation review findings (Oracle A + Oracle B) resolved
together by a single design change.
Oracle A flagged a latent value-mismatch in the HPO=on branch: the
global `self._label_horizon_candles()` call (no `pair` argument) reads
from `ft_params` only, so when `label_horizon_candles` is unset and
the helper falls back to `label_period_candles`, the logged value
reflects the global initial seed rather than the per-pair effective
value used at fit time (where HPO-tuned `label_period_candles` per
pair drives the fallback). When the user sets `label_horizon_candles`
explicitly, the global call is correct.
Oracle B flagged a cross-branch ordering inconsistency: HPO=on emitted
`label_horizon_candles` first (before per-pair lines), while HPO=off
emitted it last (after period + multiplier). Maintainers comparing
logs across HPO toggles see the same logical scalar in different
positions.
Move the per-pair effective `label_horizon_candles` into the per-pair
line in the HPO=on branch (third field after `label_period_candles`
and `label_natr_multiplier`). This closes the per-pair accuracy gap
(Oracle A) and yields the same `period → multiplier → horizon` order
in both branches (Oracle B). No behavioral change; logging only.
Jérôme Benoit [Mon, 22 Jun 2026 13:00:18 +0000 (15:00 +0200)]
refactor(quickadapter): disambiguate label_period_candles vs label_horizon_candles
Three distinct issues addressed together for terminology coherence
and operator visibility.
(1) README: replace "NATR horizon" by "NATR period" in the four rows
documenting `label_period_candles` / `min_label_period_candles` /
`max_label_period_candles` / `label_candles_step`. The noun
"horizon" was also the noun in `label_horizon_candles`, which has
the opposite temporal direction (lookahead causal-split guard vs
lookback NATR period). "NATR period" matches both the tunable name
(`label_period_candles`) and the underlying API
(`ta.NATR(..., timeperiod=label_period_candles)`).
(2) `QuickAdapterV3.set_freqai_targets`: rename the local
`label_period` (a `datetime.timedelta` spanning
`len(dataframe) * timeframe_minutes`) to `series_duration`, and
update the two log labels (`label_period: 3 days, 12:00:00` →
`series_duration: 3 days, 12:00:00`). The previous name collided
with `label_period_candles` (an int candle count) at the operator
log level; `series_duration` matches the sibling `series_length`
(an int) declared two lines below.
(3) `QuickAdapterRegressorV3.__init__` startup dump: add
`label_horizon_candles` to the "Label Parameters:" section (both
HPO-enabled and HPO-disabled branches). The parameter is a
training-time data-split lookahead (causal split guards in
`_make_train_test_split_datasets` / `_make_timeseries_split_datasets`),
NOT an HPO hyperparameter — it is neither sampled
(`trial.suggest_int`) nor bounded (`min/max_*`) nor an HPO input to
`label_objective`. Placing it in "Label Parameters:" (runtime
values) rather than "Label Hyperparameters:" (HPO config) matches
its actual semantics. In the HPO-on branch the global value is
emitted before the per-pair loop to visually distinguish it; in the
HPO-off branch it trails the existing global lines so the natural
order is `label_period_candles` → `label_natr_multiplier` →
`label_horizon_candles` (the latter falls back to the former when
unset).
The (c) clause previously named both `ValueError` and
`UnicodeDecodeError` plus their subclass relationship. The exception
classes are visible at the `except ValueError` line below — keep the
behavioral description in the docstring ("not parseable by
json.loads (malformed JSON or invalid UTF-8)") and let the code
document the catch surface.
Jérôme Benoit [Mon, 22 Jun 2026 12:35:27 +0000 (14:35 +0200)]
fix(quickadapter): catch UnicodeDecodeError in journal tail probe
Codex inline review (P2) on PR #102 flagged that
`_optuna_journal_has_corrupt_tail` catches only
`json.JSONDecodeError` from `json.loads(last_line)`, but
`json.loads` raises `UnicodeDecodeError` (subclass of
`ValueError`, NOT of `JSONDecodeError`) when the trailing record
contains invalid UTF-8 bytes — a common crash pattern when
`fsync` is interrupted mid-multibyte. The exception escapes the
helper, propagates out of `optuna_create_storage` (the helper
runs BEFORE the recoverable try/except), reaches
`optuna_create_study`'s broad outer handler, and reproduces the
silent-HPO-skip symptom under a different corruption class.
Empirically reproduced:
>>> json.loads(b'\\xc3\\x28')
UnicodeDecodeError: 'utf-8' codec can't decode byte 0xc3 in
position 0: invalid continuation byte
Broaden the except clause from `json.JSONDecodeError` to
`ValueError` — the common parent of both `JSONDecodeError` and
`UnicodeDecodeError` — so the helper treats any
`json.loads`-unparseable trailing record as corruption and routes
it through the same quarantine path. Other `ValueError` subclasses
are not plausibly raised by `json.loads` on bytes input.
Reproducer at `/tmp/quickadapter-tests/test_optuna_journal_quarantine.py`
extended from 17 to 19 scenarios (Class C4 detection + end-to-end
quarantine on invalid UTF-8). All pass.
Jérôme Benoit [Mon, 22 Jun 2026 12:31:01 +0000 (14:31 +0200)]
fix(quickadapter): quarantine corrupt optuna journal log and recover (#102)
When `JournalFileBackend`'s replay encounters a corrupt journal record, optuna raises during `JournalStorage.__init__` (Class A/B: immediate `KeyError`/`json.JSONDecodeError`) or defers the raise to the next `_sync_with_backend` (Class C1/C2/C3: truncated tail / malformed last line / bare trailing newline at EOF). Both paths previously caused HPO for the affected pair to be silently skipped on every fit cycle.
Wrap `JournalStorage` construction in a narrow `try / except (KeyError, ValueError, json.JSONDecodeError)`, atomically rename the corrupt log aside as `optuna-<COIN>.log.corrupt-<UTC_µs>`, log a `WARNING`, and retry once on a fresh path. A bounded O(1) tail probe runs BEFORE construction to detect deferred-raise EOF corruption. `OSError` is intentionally excluded so filesystem failures stay operator-actionable.
The recovery follows RocksDB's documented WAL recovery philosophy (quarantine + restart), preserves forensic evidence (rename, no `unlink`), and keeps the live-journal glob `optuna-*.log` from matching quarantined artefacts.
Jérôme Benoit [Mon, 22 Jun 2026 10:09:19 +0000 (12:09 +0200)]
refactor(quickadapter): dispatch optuna_create_sampler via match + assert_never (#103)
Follow-up to #101. Converts the `optuna_create_sampler` if/elif/else dispatch chain to a `match`/`case` statement using value patterns (`QuickAdapterRegressorV3._OPTUNA_SAMPLERS.<name>`) -- the per-field singleton `Literal[...]` typing introduced by #101 unlocks pyright/mypy exhaustiveness narrowing, so the final `case _: assert_never(sampler)` type-checks as `Never` and catches any future extension of `OptunaSampler` that forgets to add a corresponding match branch.
The `case None:` branch is structurally required (not stylistic): without it, after the four `Literal[...]` value patterns, pyright/mypy would narrow `sampler` to `None`, not `Never`, and `assert_never(sampler)` would fail to type-check. Its presence is what makes the "5th sampler addition -> type error at assert_never" claim effective. The error message inside `case None:` is preserved verbatim from the prior `else: raise ValueError(...)` for user-facing wire compatibility.
Behavior delta is confined to non-`Literal` string inputs (e.g. dynamically-injected `"garbage"`): prior code raised `ValueError` with the supported-values list; new code raises `AssertionError` with the standard `typing.assert_never` message. In practice this path is unreachable from every current in-repo call site -- the sole caller `optuna_create_study` already validates `sampler not in samplers` and raises `ValueError` with the supported-values list before dispatching, so misconfigured `_optuna_config["sampler"]` values surface the old error shape from the upstream gate. The new `AssertionError` would only manifest via a hypothetical future caller that bypasses `optuna_create_study`, which does not exist today.
Pattern parity: the same `assert_never` exhaustiveness idiom is already used in this file for `support_policy` dispatch; value-pattern syntax matches the 8 call sites migrated by #101.
Closes the deferred follow-up from PR #101 (issue #88).
Jérôme Benoit [Mon, 22 Jun 2026 09:32:44 +0000 (11:32 +0200)]
refactor(quickadapter): consolidate Optuna sampler tuples to NamedTuple (#101)
PR #81 consolidated `_OPTUNA_NAMESPACES` to `Utils._OptunaNamespaces` (a `NamedTuple` with per-field singleton `Literal` types). Propagate the same pattern to the three sibling sampler tuples in `QuickAdapterRegressorV3.py`:
Each new `_Optuna*Samplers` NamedTuple class lives at module level immediately before `class QuickAdapterRegressorV3`, matching the `_OptunaNamespaces` adjacency in `Utils.py`. Per-field types are singleton `Literal["..."]` (not the `OptunaSampler` union) to unlock pyright/mypy narrowing for a future `assert_never` migration. The class-private `_OPTUNA_*_SAMPLERS` instance constants remain class-attributes of `QuickAdapterRegressorV3` to preserve every consumer's `QuickAdapterRegressorV3._OPTUNA_*` access pattern.
Each `_Optuna*Samplers` redeclares its field defaults (per-class `Literal["tpe"] = "tpe"` etc.), replacing the prior tuple-slice derivation `_OPTUNA_HPO_SAMPLERS = _OPTUNA_SAMPLERS[:2]` and the 6-line custom reordering of `_OPTUNA_LABEL_SAMPLERS`. Trade-off: minor string-literal duplication across 3 classes is accepted for harmonization with the `_OptunaNamespaces` template; the single source of truth for the 4 valid tokens stays the `OptunaSampler = Literal[...]` alias unchanged at module level.
Migrates 8 positional-indexing call sites (`_OPTUNA_*_SAMPLERS[N] # "name"`) to attribute access (`.<name>`); drops the 6 surviving inline annotations at the call sites (2 sites carry no annotation pre-refactor; the multi-line site collapses 3 source lines into 1) plus the 4 analogous annotations inside the deleted 6-line `_OPTUNA_LABEL_SAMPLERS` custom-ordering block. The `_OPTUNA_HPO_SAMPLERS_SET` and `_OPTUNA_LABEL_SAMPLERS_SET` frozenset companions are kept unchanged (still used in O(1) membership testing at `optuna_samplers_by_namespace`); their type annotation `Final[frozenset[OptunaSampler]]` is preserved.
Non-migration sites confirmed unchanged: the frozenset companion construction, the `', '.join` error-message iteration over `_OPTUNA_SAMPLERS`, and the `_SET` membership references. All iterate over the NamedTuple instance and produce byte-identical output.
Add `NamedTuple` to the existing `from typing import (...)` block (alphabetical, between `Literal` and `Optional`). `assert_never` already imported, kept in anticipation of the deferred follow-up.
Per-field singleton `Literal[...]` unlocks `assert_never` exhaustiveness on the `optuna_create_sampler` dispatch chain -- left to a follow-up PR since the migration (`else: raise ValueError(...)` becoming `assert_never(sampler)`) changes the user-facing error contract on the unreachable branch.
The AGENTS.md *Canonical defaults* principle and the README documented enum order are encoded in the field-declaration order of each `_Optuna*Samplers`. NamedTuple remains a `tuple` subclass, so `[0]` indexing, `len(...)`, `frozenset(...)`, `', '.join(...)`, and iteration all keep their existing semantics. No behavior change.
Reviewed by two pre-implementation design passes (3-oracle on v1; Metis + Momus + meta-Oracle on v2) and a 3-oracle review on the live PR, each citing upstream evidence from `freqtrade/freqai/` confirming no external consumer.
Follow-up from PR #81 review (Oracle harmonization dimension).
Prose state-form (`Utils.py`, `QuickAdapterRegressorV3.py`):
- `_normalize_label_column_name` docstring: ``Raises ValueError when
the result contains `&` or `%` after sigil strip``.
- Deprecated-config-key warning aligns with the sibling pattern at
the adjacent branch: ``f"{old_path!r} is deprecated, use
{new_path!r} instead"``.
- `sanitize_and_renormalize` docstring states ``mean(out) == 1`` as the
rescale invariant.
- Optuna-label throttle log reads ``callback throttled,
{N} candles until next emission``.
- Fit-live-predictions warmup log reads ``Fit live predictions not
warmed up: {N} candles until warmup completion``.
Docstrings on validator/composer helpers (3 functions lacking a
docstring at HEAD):
- `_apply_support_policy`: documents the ``policy='raise'`` /
``policy='fallback'`` dispatch contract.
- `_compose_train_weights_with_support`: documents the support-gating
flow (None-label-weights branch routes through
``_apply_support_policy`` when ``strategy != 'none'``; main branch
composes and validates the summary against three thresholds).
- `_validate_optuna_label_best_params`: enumerates the rejection
paths and the optional ``expected_selection_metadata`` drift gate.
Harmonization (post-merge carry-over):
- `LABEL_WEIGHT_SUFFIX` renamed to `_LABEL_WEIGHT_SUFFIX`
(no external consumer; symmetric with
`_LABEL_KNOWN_AT_LOOKAHEAD_SUFFIX`).
- `safe_distribution_fit` call-site contexts harmonized with the
PR #97 / PR #99 ``[<pair>] <event>`` convention:
`f"[{pair}] di_values_weibull_fit"` and
`f"[{pair}] label_norm_fit:{label_col}"`.
Python idioms:
- `_adapt_label_generator` rejects any 3rd positional parameter
whose name is not ``logger``, regardless of whether the parameter
is required or has a default. A defaulted non-``logger`` 3rd
positional raises ``ValueError`` at registration. The 3-arg
pass-through is reached only when ``positional[2].name == "logger"``.
- `_build_sample_weight_inputs` switches the two `logger.debug`
calls to lazy ``%s`` formatting so the f-string body is not
materialized when the debug level is disabled.
Consolidates SAFE + DRY + COSMETIC findings from the 4-axis review of
`add1fb7..d6a718f` (PRs #90, #94, #95, #96, #97 + 2 style + d6a718f).
Migration (PR #94 carry-over to `QuickAdapterV3.py`):
- `_TRADE_DIRECTIONS_SET` and `_ORDER_TYPES_SET` are
`Final[frozenset[T]]` constants adjacent to the canonical
`_TRADE_DIRECTIONS` / `_ORDER_TYPES` tuples; the strategy reads
them directly at all 3 call sites. The
`@staticmethod @lru_cache(maxsize=None)` set-views are absent at
HEAD.
Context-prefix harmonization (PR #97 carry-over):
- The 2 `sanitize_and_renormalize` calls in
`QuickAdapterRegressorV3._apply_pipelines` carry the `[<pair>]`
prefix (`f"[{pair}] post_feature_pipeline:train"` / `:test`),
matching the PR #97 convention at the 4 `compose_sample_weights`
call sites.
Alias removal (PR #96 carry-over):
- `Utils.label_known_at_column_name` is absent at HEAD. The
underlying column has zero external callers in the repo, and
PR #96 shifted semantics (absolute index -> per-row offset);
no back-compat alias is warranted.
DRY (`d6a718f` carry-over):
- The 8-tuple swap blocks in `_make_train_test_split_datasets`
and `_make_timeseries_split_datasets` use pythonic parallel
pair-swap (`a, b = b, a` per slot pair); 4 lines per slot pair
instead of the 18-line 8-tuple parallel assignment.
Dead constant removal:
- `QuickAdapterRegressorV3._AGGREGATE_DISTANCE_METRICS_SET` is
absent at HEAD; the constant has no caller. The
`_CLUSTER_DENSITY_DISTANCE_METRICS_SET` block comment lists the
7 aggregate metrics inline for reference (`harmonic_mean`,
`geometric_mean`, `arithmetic_mean`, `quadratic_mean`,
`cubic_mean`, `power_mean`, `weighted_sum`).
Cosmetic:
- `_LABEL_KNOWN_AT_LOOKAHEAD_SUFFIX` placement is adjacent to
`LABEL_WEIGHT_SUFFIX` (both are label-aux column suffixes).
- `_known_at_lookahead` returns `int64` on both single-series and
multi-series paths (symmetric dtype contract).
- Schema-version reset log: `f"v{existing_schema_version!r}"`
renders unambiguously for non-int corrupt values (booleans,
strings).
Jérôme Benoit [Mon, 22 Jun 2026 01:54:59 +0000 (03:54 +0200)]
fix(quickadapter): route reversed train weights through support_policy (#98)
PR #85 added `_compose_train_weights_with_support` (gates training-set weights through `support_policy`) and `_compose_eval_weights` (eval-side, deliberately bypasses `support_policy`). The `reverse_train_test_order` path in `_make_train_test_split_datasets` and `_make_timeseries_split_datasets` swapped slices AT THE FINAL `build_data_dictionary` call -- AFTER weight composition -- so the actual-train slot received weights composed by `_compose_eval_weights` (silent bypass), and the actual-test slot received weights composed by `_compose_train_weights_with_support` (wrong direction, typically a no-op under `support_policy='fallback'` default).
Reachable only under `causal_mode=false` (deprecated; acausal baselines only) since `causal_mode=true` rejects `reverse_train_test_order=true` upfront.
Fix: perform the train/test slice swap BEFORE weight composition so the `train_*` and `test_*` identifiers map to their actual training roles throughout. Both call sites converge to a single `dk.build_data_dictionary` return; context strings in `support_policy` log/raise messages now reflect the true train/test role.
Add an upfront `ValueError` in `_make_train_test_split_datasets` when `test_size=0` AND `reverse_train_test_order=True`, mirroring the existing `causal_mode`/reverse rejection pattern. The `timeseries_split` path already rejects `test_size < 1` upstream of the swap.
Behavior change in the deprecated path: `support_policy='raise'` now correctly raises on actual-train insufficient support; `support_policy='fallback'` now correctly warns.
Reviewed by three parallel Oracle passes (math + algorithmics + scope/reachability; Python state-of-the-art + harmonization + implementation elegance; documentation + terminology + completeness) at design stage and again post-implementation, each citing upstream evidence from `freqtrade/freqai/`.
Follow-up from PR #80 review, deferred during PR #90.
Jérôme Benoit [Mon, 22 Jun 2026 01:34:16 +0000 (03:34 +0200)]
fix(quickadapter): prefix `compose_sample_weights` and `sanitize_and_renormalize` log entries with caller context (#97)
`compose_sample_weights` and `sanitize_and_renormalize` accept a
required keyword-only `context: str` parameter; every warning, error,
and inner `sanitize_and_renormalize` call uses `context` as its sole
prefix.
- `compose_sample_weights(..., *, logger, context, on_collapse=...)`:
`context: str` keyword-only, required. The 4 internal
warnings/errors (shape-mismatch `ValueError`, all-dropped
`LabelWeightSupportError`, sparse-mass warning, collapse-on-
survivors `LabelWeightSupportError` and fallback warning) prefix
with `{context}:`. Internal `sanitize_and_renormalize` calls
compose sub-contexts: `{context}:base_only`,
`{context}:label_weighted`, `{context}:base_fallback`. The sparse-
mass message reads `sparse weighting mass`, accurate for both
train and eval paths.
- `sanitize_and_renormalize(..., *, logger=None, context: str)`:
`context` keyword-only, required. The 5 warnings/errors
(`drop_mask` shape `ValueError`, `drop_mask` dtype `ValueError`,
rescale-overflow warning, weights-collapsed warning,
mask-covers-all warning) prefix with `{context}:`. The redundant
`(context=%s, ...)` subfield is dropped.
- 5 `compose_sample_weights` call sites in
`QuickAdapterRegressorV3.py` pass `context=context`. The 2
external `sanitize_and_renormalize` call sites
(`post_feature_pipeline:train`, `post_feature_pipeline:test`)
already pass `context=`.
Log format goes from
`compose_sample_weights: sparse training mass (59/2603 rows ...)`
to
`[ETH/USDT] train_test_split:train: sparse weighting mass (59/2603 rows ...)`.
The pair, split method, and train/eval side are traceable in the
log line.
Jérôme Benoit [Mon, 22 Jun 2026 01:21:28 +0000 (03:21 +0200)]
refactor(quickadapter): migrate _*_set() lru_cache family to Final[frozenset] (#94)
Each `@staticmethod @lru_cache(maxsize=None) _X_set()` family member
maps to a class-level `_X_SET: Final[frozenset[...]]` adjacent to its
backing tuple. Derived members express their content through set
algebra over the new `_*_SET` constants.
Call sites at `(cls|QuickAdapterRegressorV3)._X_set()` resolve to the
matching `_X_SET` reference. `optuna_samplers_by_namespace` returns
`tuple[frozenset[OptunaSampler], OptunaSampler]` to match the new
constant types. The surviving `@lru_cache(maxsize=8)` on
`optuna_samplers_by_namespace` falls outside the `_*_set` family and
stays in place.
Jérôme Benoit [Mon, 22 Jun 2026 01:20:11 +0000 (03:20 +0200)]
refactor(quickadapter): rename known_at_index to known_at_lookahead (#96)
PR #95 retained the historical column name `<label>_known_at_index` for what is now a per-row label lookahead in candles, to keep that hotfix strictly minimal. This PR converges the column suffix, the helper, the dataclass field, the static method, and the per-call-site locals onto `_known_at_lookahead`, with a retro-compat alias on the only externally-named public helper (`label_known_at_column_name = label_known_at_lookahead_column_name`).
The auxiliary `<label>_known_at_*` column is regenerated on every training run inside `set_freqai_targets`; FreqAI persists only the fitted model and `extra_returns_per_train`, never auxiliary dataframe columns -- the rename invalidates no on-disk artifact.
Reviewed by three parallel Oracle passes (math + claims-coherence; Python state-of-the-art + harmonization; documentation + terminology + PR-description coherence), each citing upstream evidence from `freqtrade/freqai/freqai_interface.py`, `data_kitchen.py`, and `data_drawer.py`. Consensus fixes were applied: README `causal_mode` formula symbol bound to the column token (`row-wise max(<label>_known_at_lookahead)`) to colocate definition with usage.
The two causal-guard local variable pairs were also harmonized to the local `train_<noun>` family (`train_known_at_lookahead`, `train_known_at_position`) used by the surrounding `_make_*_datasets` methods.
Jérôme Benoit [Mon, 22 Jun 2026 01:02:56 +0000 (03:02 +0200)]
fix(quickadapter): use slice-invariant lookahead for causal guard (#95)
* fix(quickadapter): use slice-invariant lookahead for causal guard
PR #78 stored '<label>_known_at_index' as 'arange(len) + horizon +
kernel_half_width' -- absolute positions in the dataframe passed to
'set_freqai_targets'. freqtrade's 'dk.slice_dataframe' (a '.loc' filter)
runs AFTER 'set_freqai_targets' and drops warmup rows but preserves
column values, so those pre-slice positions survived into the post-slice
'unfiltered_df'. The causal guard then compared them against
'first_test_position' derived from 'np.arange(len(unfiltered_df))' --
local post-slice positions in a different coordinate system. The unit
mismatch wiped out most or all training rows on every pair.
Production crash on 2026-06-22 (XRP/USD): "removed 2621
causal-unsafe train rows" followed by "causal guard removed all
train rows, skipping".
Fix: the column now stores a per-row label lookahead (in candles),
invariant under 'dk.slice_dataframe'. Consumers combine the row's
local position with the lookahead to recover the local known-at
position before comparing to 'first_test_position'. Column name
'<label>_known_at_index' is retained for this hotfix; a rename to
'<label>_known_at_lookahead' (with rétro-compatible alias) is left
to a follow-up PR per AGENTS.md 'small, verifiable changes'.
Touches:
- Utils.py: producer rewritten to store a constant per-row lookahead;
'LabelData' and 'label_known_at_column_name' docstrings document the
new contract; '_LABEL_KNOWN_AT_SUFFIX' carries an inline disambiguation.
- QuickAdapterV3.py: smoothing-lookahead advance comment harmonized to
the canonical 'per-row label lookahead (in candles)' phrasing.
- QuickAdapterRegressorV3.py: '_known_at_index' docstring rewritten;
'train_test_split' and 'timeseries_split' causal-mode branches add
'train_positions + delta' before the '< first_test_position' check;
'timeseries_split' hoists 'train_positions' for symmetry with
'train_test_split'.
- README.md: 'causal_mode' tunable description reflects the new
comparison semantic.
Reviewed by three parallel Oracle passes (math/algo/scope,
Python state-of-the-art / harmonization, documentation /
terminology / concision) with cross-validation; one false alarm
on a missing position-only fallback in 'timeseries_split' was
resolved by confirming 'TimeSeriesSplit.gap' enforces the
chronological purge at the sklearn layer.
* docs(quickadapter): shrink _known_at_index docstring to LabelData pointer
Per multi-oracle PR #95 review (Oracle 3 §8.1): paragraph 1 of
_known_at_index duplicated the slice-invariance rationale already
canonical on LabelData.known_at_index. Replace with a thin pointer per
AGENTS.md *No duplication: maintain single authoritative documentation
source; reference other sources rather than copying.*
Consolidates P1/P2 findings from `chatgpt-codex-connector` review
comments on PRs #78, #79, #80, #81, and PR #90.
Utils.py + label generation:
- `_generate_extrema_label` accepts `logger: Logger | None`; the
`LabelGenerator` type signature, `generate_label_data` dispatcher,
and `QuickAdapterV3.set_freqai_targets` caller propagate the logger.
`_generate_extrema_label` has no `F821 logger` undefined-name path.
- `register_label_generator` routes the input through
`_adapt_label_generator`. The adapter detects the canonical
`(dataframe, params, logger)` shape by a positional parameter named
`logger` at index 2 (with or without a default); other generators
with 2 required positional parameters are wrapped via
`functools.wraps` (preserves `__name__`/`__doc__`/`__wrapped__`) to
drop the logger argument at dispatch, with defaulted positionals
after index 1 left at their defaults. `ValueError` is raised at
registration for `*args`, `**kwargs`, keyword-only `logger`, fewer
than 2 required positionals, more than 3 required positionals, and 3
required positionals whose third name is not `logger`.
- `safe_divide` denominator zero-check uses exact-zero
(`denominator_arr != 0.0`); subnormal and satoshi-scale denominators
pass the gate, and non-finite division outputs coerce to `fallback`
via the post-division finite-mask.
Causal label split lookahead:
- `QuickAdapterV3.get_label_horizon_candles` recomputes the horizon
from the current `label_period_candles` (via
`get_label_period_candles`); init omits `label_horizon_candles`, so
HPO period updates propagate to the horizon. The regressor's
`_optuna_label_params` init likewise omits `label_horizon_candles`.
- `QuickAdapterV3.set_freqai_targets` advances `<label>_known_at_index`
by the smoothing kernel half-width after smoothing. The
`Utils.get_smoothing_kernel_half_width(config, *, series_length)`
helper reuses `get_odd_window`/`get_even_window`/`get_savgol_params`
shared with `smooth()` and dispatches on kernel routing:
- `filtfilt`-routed zero-phase kernels (members of
`SMOOTHING_KERNELS`: `gaussian`, `kaiser`, `kaiser_bessel_derived`,
`triang`): half-width `effective_window - 1` (forward+backward
pass extends dependency to the full filter length on both sides).
- Single-pass centered windows (`smm`, `sma`, `savgol`): half-width
`effective_window // 2`.
- `gaussian_filter1d`: `int(4.0 * sigma + 0.5)` matching
`scipy.ndimage` default truncation.
The helper returns 0 when `smooth()` itself short-circuits:
`series_length < max(window_candles, 3)` (top-level no-op) and, for
the filtfilt/savgol routes, `series_length < effective_window`
(downstream short-series no-op in `zero_phase_filter` /
`savgol_filter`). The `series_length` parameter is keyword-only and
required.
Label-weighting support policy:
- `QuickAdapterRegressorV3._compose_train_weights_with_support` routes
the zero-pivot case (label-weighting strategy configured but no
label weights available) through `_apply_support_policy`; the
support policy governs the contract: `raise` raises, `fallback`
warns.
Label Optuna selection hardening:
- `_OPTUNA_LABEL_SELECTION_SCHEMA_VERSION` is `2`, co-located with
`_OPTUNA_LABEL_BEST_PARAMS_SCHEMA_VERSION` in `Utils`. The two
constants are independent: wire format and selection algorithm
carry separate version axes.
- `_optuna_label_selection_metadata` rejects non-finite `label_weights`
/ `label_p_order` with `ValueError`; downstream dict equality on the
selection_metadata is NaN-safe.
- `_validate_optuna_label_best_params` accepts an
`expected_selection_metadata` keyword. It rejects files that are
not a `{schema_version, params, selection_metadata}` dict, files
with mismatched `schema_version`, files missing or with an invalid
`selection_metadata.schema_version`, and -- when
`expected_selection_metadata` is provided -- files whose stored
`selection_metadata` differs from the caller's current view. Legacy
unversioned best-params files (no `schema_version`) are rejected
outright. `QuickAdapterRegressorV3.optuna_load_best_params` passes
its `_optuna_label_selection_metadata()` view for the label
namespace; `QuickAdapterV3.optuna_load_best_params` omits the
keyword, since the strategy reads only `label_period_candles`,
`label_horizon_candles`, and `label_natr_multiplier` from the
best-params file.
- The study schema-migration branch resets the Optuna label study
whenever `existing_schema_version` is not an `int` (rejects `bool`
and other types) or differs from `target_version`; an unversioned
study (legacy pre-metadata) and a version-mismatched study are
treated identically. Trials selected under a different metric
whitelist or weighting scheme cannot be reused.
- `_optuna_label_selection_metadata` includes `label_weights` and
`label_p_order`; `_calculate_distances` consumes both, so the
idempotent `set_user_attr` write detects drift on those tunables.
- `_calculate_distances` validates `label_weights` length against the
original objective count (raises on mismatch) only on the slicing
path (when `objective_indices is not None` and the count differs).
The sliced vector falls back to uniform weights only when
`np.all(sliced == 0.0)` (the user's positive weights all on
dropped objectives); slices containing negative or non-finite
values flow through to `_validate_label_weights`.
- `_validate_label_selection_metric` accepts an `aggregate_allowed`
parameter; cluster/density category callers pass `False`,
restricting the valid set via the cached
`_cluster_density_distance_metrics_set()` classmethod
(SciPy-compatible non-probability metrics). The aggregate metrics
(`harmonic_mean`, `geometric_mean`, `arithmetic_mean`,
`quadratic_mean`, `cubic_mean`, `power_mean`, `weighted_sum`) live
in the cached `_aggregate_distance_metrics_set()` classmethod
derived by set-algebra from
`_distance_metrics_set() - _scipy_metrics_set() -
_probability_distance_metrics_set()`. README cluster/density metric
rows list the SciPy-compatible set. `compromise_programming`/
`topsis` accepts the full aggregate set via
`_calculate_trial_distance_to_ideal`.
Out of scope:
- Two PR #80 findings: `reverse_train_test_order` support-policy
routing; post-`feature_pipeline.fit_transform` support recheck.
Jérôme Benoit [Sun, 21 Jun 2026 21:45:02 +0000 (23:45 +0200)]
fix(quickadapter): stabilize label optuna selection (#81)
Multi-objective `label` Pareto best-trial selection in the
QuickAdapter regressor.
- Probability-style metrics (`jensenshannon`, `hellinger`,
`shellinger`) rejected for `label_distance_metric` /
`label_cluster_metric` / `label_density_metric` — Pareto objective
matrices are unbounded floats, not probability vectors. Invalid
values fall back to `euclidean` with a warning.
- Constant objective dimensions dropped before trial-distance
computation: a constant objective dimension is non-informative and
would bias the geometry. Tolerance at
`_NON_CONSTANT_OBJECTIVE_ATOL: Final[float] = 1e-8`.
- User-supplied `label_weights` matching the original objective count
slice to align with the non-constant subset; mismatched sizes flow
to `_validate_label_weights(mode="raise")`.
- Deterministic best-trial tie-break by `(distance, trial.number)`,
independent of `study.best_trials` ordering.
- All-constant Pareto front falls back to the lowest `trial.number`.
- Persisted Optuna `label` study user-attr `selection_metadata` nests
`schema_version`
(`_OPTUNA_LABEL_SELECTION_SCHEMA_VERSION: Final[int] = 1`) and
`method_config`. Studies without a recorded `schema_version` are
tagged at the current version on next `optuna_create_study` (trials
preserved); studies recording a different version are reset. The
`selection_metadata` write is idempotent: skipped when unchanged,
warned on diff.
- `OptunaNamespace` Literal and `_OPTUNA_NAMESPACES` (a `NamedTuple`
of `hp` and `label` with per-field literal types) live in `Utils`,
with `.hp` / `.label` accessors at all call sites
(`QuickAdapterRegressorV3`, `QuickAdapterV3`). No per-class tuple
alias, no inline `# "hp"` / `# "label"` annotations.
Jérôme Benoit [Sun, 21 Jun 2026 20:37:10 +0000 (22:37 +0200)]
fix(quickadapter): restore _format_collection ordering before register decorators
Rebase of the label-weight-support-policy PR onto main (atop the style line-wrap commit) reapplied the _format_collection hunk at the wrong location, placing the helper AFTER its @_format_value.register(list|tuple|set|dict|np.ndarray) decorators. Move the definition back to its canonical position above the dispatchers.
Jérôme Benoit [Sun, 21 Jun 2026 20:26:45 +0000 (22:26 +0200)]
feat(quickadapter): add label weight support policy (#85)
Per-row sample-weight composition decoupled from the final per-split
compose. Causal-guard filtering operates on raw base/label weights.
Configurable support thresholds gate the train-weight compose.
- `SampleWeightInputs` dataclass carries `(base, label,
label_weighting_config)` from `_build_sample_weight_inputs`;
`compose_sample_weights(base, label)` runs AFTER
`train_test_split`/`TimeSeriesSplit` AND AFTER causal guards.
`__post_init__` validates 1-D shape, base/label shape parity,
required `label_weighting_config` keys, and `support_policy` enum
membership.
- Thresholds in `freqai.label_weighting`:
`min_pivot_equivalent_count` (default 3),
`min_positive_label_weight_fraction` (default 0.01),
`min_effective_sample_size` (default 3.0; Kish ESS).
- `support_policy: enum {fallback, raise}` (default `fallback`)
drives the failure mode when any threshold trips. Eval (test/val)
weights bypass this policy by design.
- `compose_sample_weights` takes `on_collapse: Literal["raise",
"fallback"]` (default `"raise"`); the train path lets collapse raise
through `LabelWeightSupportError` so `support_policy` catches it,
the eval path passes `"fallback"` to preserve the label-derived
drop mask.
- `_apply_support_policy` (typed `policy: LabelWeightSupportPolicy`)
dispatches `fallback` and `raise` branches via `match`/`case` with
`assert_never` exhaustiveness; `compose_sample_weights` applies the
same pattern on `on_collapse`.
- `_shuffle_split_rows` 4-tuple shuffler covers label weights.
- `_filter_train_by_mask` accepts optional `train_label_weights` for
uniform causal-guard filtering across base+label weight arrays.
- `LabelWeightSupportSummary`, `_effective_sample_size` (Kish ESS),
and `summarize_label_weight_support` documented as operational spec.
README documents the four `label_weighting` rows;
`config-template.json` records the `fallback` default.
Shared finite-sample, guarded distribution-fit, safe divide/log-ratio,
and sigmoid-domain helpers. Log/division feature paths route through
the helpers; distribution fits guard empty, non-finite, and constant
samples.
- `Utils.py` helpers: `FiniteSample` dataclass with `finite_sample`;
`safe_distribution_fit` (documented fallback-length contract);
`safe_divide`; `safe_log_ratio`.
- `nan_average` finite/zero-weight guards; documented divergence from
`np.nanmean` (strips +/-inf as well as NaN; bounded for current
callers).
- `_clip_sigmoid_domain` in `LabelTransformer.py` guards
`sp.special.logit` against values outside the open `(-1, 1)` domain
during `sigmoid` inverse normalization.
- `feature_engineering_expand_basic` and Utils log/divide sites
(`top_log_return`, `bottom_log_return`, `price_retracement_percent`,
`ewo` normalize, `zigzag` log prices, KC/BB/VWAP widths) route
through the safe helpers.
- DI Weibull and label `norm` fits in `fit_live_predictions` use
`safe_distribution_fit`; DI cutoff fallback at
`_DI_CUTOFF_DEFAULT: Final[float] = 2.0`.
Jérôme Benoit [Sun, 21 Jun 2026 18:01:23 +0000 (20:01 +0200)]
feat(quickadapter)!: add causal label split foundation (#78)
Causal split guards on QuickAdapter training. Default causal mode
rejects `data_split_parameters.shuffle=true`,
`feature_parameters.shuffle_after_split=true`, and
`feature_parameters.reverse_train_test_order=true`.
- `feature_parameters.causal_mode` (default `true`): guard toggle.
`false` is deprecated.
- `feature_parameters.label_horizon_candles` (default
`label_period_candles`): candles after a label row before its label
is considered known by causal split guards. Fallback chain
`label_horizon_candles` -> `label_period_candles` -> `1`.
- `<label>_known_at_index` columns expose `LabelData.known_at_index`
per-row; multi-label boundary via element-wise max across present
columns.
- `timeseries_split` `gap` auto-set from `label_horizon_candles` under
causal mode; explicit `gap < label_horizon_candles` rejected.
- Persisted Optuna `label` best-params JSON has shape
`{schema_version, params}`
(`_OPTUNA_LABEL_BEST_PARAMS_SCHEMA_VERSION = 2`). Unversioned files
identified by shape; version-mismatched files emit distinct
"missing" vs "incompatible" warnings.
- `_label_aux_column_name` shared sigil-stripping helper backs
`label_weight_column_name` and `label_known_at_column_name`;
uniform collision guard against `&`/`%` and empty stem.
- `QuickAdapterRegressorV3.version = 3.12.0`.
BREAKING CHANGE: `feature_parameters.causal_mode` defaults to `true`.
Configs with `data_split_parameters.shuffle=true`,
`feature_parameters.shuffle_after_split=true`, or
`feature_parameters.reverse_train_test_order=true` raise at training
time.
Add a fourth off-pivot weighting mode that superposes the epsilon
floor and the gaussian bumps additively, and fix a related defect in
_scatter_weights that allowed pivot rows to sit *below* the off-pivot
field whenever that field at the pivot's index could legitimately
exceed the pivot's raw weight.
with phi = eps * B(W) (B mean or median, eps in [0, 1]; phi = 0 on
empty pivots or non-finite baseline), and per-pivot sigma_p from
_compute_pivot_sigmas (fixed or k-NN). The combined formulation reuses
both existing closed forms verbatim:
Bound: phi <= f(i) <= phi + max_p w_p. The new mode reduces to pure
gaussian when eps = 0 (bit-identical). The reduction is a per-row max
over per-pivot Gaussian bumps; phi is the epsilon floor.
Before this change _scatter_weights wrote out[p] = w_p unconditionally,
so a pivot whose raw weight was below the off-pivot field at its
index appeared as a sharp dip relative to its neighbors. Two
manifestations of the same defect class:
- 'epsilon' / 'epsilon_gaussian' (sub-floor): a pivot with
w_p < phi (e.g. W = (0.001, 1.0, 1.0) with eps = 0.5 and
B = median, phi = 0.5) sat at 0.001 while neighbor rows sat at phi.
- 'gaussian' / 'epsilon_gaussian' (sub-bump): a weak pivot with a
strong neighbor (e.g. W = (0.001, 1.0) at indices (0, 1), sigma = 1)
sat at 0.001 while the off-pivot field at the pivot's own index was
1.0 * exp(-0.5) ~= 0.6065 (the neighbor's gaussian bump).
Both cases are corrected by a single uniform change: _scatter_weights
now writes out[p] = max(w_p, fill[p]) so pivot rows are never written
below the off-pivot field. 'zero' is bit-identical (fill is always 0,
so max(w_p, 0) = w_p when w_p >= 0). 'gaussian' in the sparse-pivot
regime (the typical configuration, especially with k-NN bandwidth) is
also bit-identical because fill[p] equals w_p when no neighbor's bump
at p exceeds w_p.
Implementation
--------------
- _scatter_weights: pivot rows take np.maximum(weights, fill_weights)
unconditionally. Off-pivot rows unchanged.
- _compute_epsilon_floor (renamed from _epsilon_floor): extracted
helper that returns phi (mean / median / fallback). Reused by
'epsilon' and 'epsilon_gaussian'. Parameter baseline narrowed to
the FillEpsilonBaseline Literal type.
- _compute_gaussian_bumps (renamed from _gaussian_bumps): extracted
adapter over _gaussian_fill_weights. Reused by 'gaussian' and
'epsilon_gaussian'. logger is kwarg-only.
- compute_label_weights: dispatcher gains the FILL_METHODS[3] branch.
The combined branch computes bumps once and adds phi in-place via
np.add(out=fill_weights), keeping peak memory at the existing
(chunk, M) buffer; phi is constant in p so the post-reduction add
is algebraically identical to adding inside the chunk loop while
saving O(chunk * M) writes. ValueError messages tightened to
include 'supported values are ...' for parity with
_compute_pivot_sigmas and _aggregate_metrics.
- LabelTransformer.py: extends FillMethod Literal and FILL_METHODS
tuple with 'epsilon_gaussian' at index 3. No new tunables, no new
validators (the existing _EnumValidator(FILL_METHODS) picks up the
new value automatically; existing range / type validators on
fill_epsilon / fill_sigma_* / fill_bandwidth_* apply unchanged).
- QuickAdapterV3.py: logging block refactored from if/elif chain to
parallel if blocks keyed on tuple membership so epsilon and sigma
parameter groups emit independently for each mode that uses them.
Documentation
-------------
README cells updated with set-membership 'Ignored when ...' clauses
matching the new index sets (epsilon | epsilon_gaussian for the
floor parameters, gaussian | epsilon_gaussian for the kernel
parameters). The fill_method description names the additive
composition explicitly and the pivot-row lift invariant
(out[p] = max(w_p, f(p))).
Verified manually on the host via AST extraction harness (no automated
test infrastructure exists in quickadapter/):
- zero mode: bit-exact with prior code (fill is 0, max(w_p, 0) = w_p).
- gaussian mode, sparse pivots: bit-identical to prior code (no
neighbor's bump at p exceeds w_p, so the lift is a no-op).
- gaussian mode, neighbor-dominated regime: pivot rows lifted to the
local field max, fixing the sub-bump dip. Verified with the
counterexample W = (0.001, 1.0) at indices (0, 1), sigma = 1:
legacy out[0] = 0.001, fixed out[0] = 1.0 * exp(-0.5) ~= 0.6065.
- epsilon back-compat (above-floor pivots): phi = eps * mean(W)
reproduced; pivots above phi unchanged.
- epsilon pivot-dip fix: W = (0.001, 1.0, 1.0), eps = 0.5,
baseline = median; legacy out[0] = 0.001, fixed out[0] = phi = 0.5.
- epsilon_gaussian with eps = 0: bit-identical to pure gaussian.
- epsilon_gaussian additive decomposition: out_eg - out_g = phi at
every off-pivot row.
- epsilon_gaussian pivot-row lifted: W = (0.001, 1.0, 1.0) at
well-separated indices (e.g. (0, 100, 200)), eps = 0.5,
baseline = median, sigma = 2.0; out[0] = phi + 0.001 ~= 0.501
(was 0.001 before the scatter fix).
- empty pivots: all four modes return all-zero.
- negative pivot weights still rejected by _gaussian_fill_weights.
- knn bandwidth + epsilon_gaussian: finite, bounded below by phi.
- ValueError messages on invalid fill_method / fill_epsilon_baseline
include 'supported values are ...'.
Jérôme Benoit [Thu, 4 Jun 2026 22:29:58 +0000 (00:29 +0200)]
feat(label_weighting): adaptive k-NN bandwidth for gaussian off-pivot fill (#77)
* feat(label_weighting): adaptive k-NN bandwidth for gaussian off-pivot fill
Address the crushing of weaker pivots by stronger neighbors when pivots
fall within ~sigma_candles of each other in fill_method='gaussian'. The
per-row max aggregator preserves the upper bound Out[i] <= max_p w_p
but a wide constant sigma lets a strong neighbor's Gaussian dominate a
weak pivot's tail.
Add a k-nearest-neighbor bandwidth selector (Loftsgaarden &
Quesenberry 1965; Silverman 1986, paragraph 5.2) that adapts each
pivot's sigma to local pivot density:
where d_k(p) is the index distance to the k-th pivot neighbor. The
upper bound on Out[i] is preserved (no over-amplification) and dense
clusters automatically contract their Gaussians to stop overlapping.
Implementation:
- Pivots are emitted chronologically by zigzag, so the 1D k-NN reduces
to a sliding k-window over sorted indices, O(M) without a spatial
index.
- _gaussian_fill_weights accepts a per-pivot sigma vector via NumPy
broadcasting; the existing chunked exp/multiply/max kernel is
unchanged.
- Default fill_bandwidth='fixed' preserves byte-for-byte the previous
algorithm.
Adds "uniform" to WEIGHT_STRATEGIES (between "none" and the metric
names) which assigns weight=1.0 to every detected pivot. Off-pivot rows
remain governed by the existing fill_method (zero / epsilon / gaussian),
so uniform + gaussian collapses cleanly to a pure proximity kernel
around each pivot. Naming follows sklearn convention
(KNeighbors(weights="uniform"), DummyClassifier(strategy="uniform")).
* chore(quickadapter): bump strategy and regressor version 3.11.11 -> 3.11.12
* refactor(weights): hoist indices_array and valid_mask in compute_label_weights
Compute indices_array and valid_mask once at the top of the function
instead of after the strategy dispatch. The uniform branch can now use
indices_array.size instead of len(indices), and the duplicate
np.asarray / valid_mask construction lower in the function is removed.
Saves one np.asarray and one mask computation per call.
Drop the redundant indices: list[int] parameter now that the only
caller (compute_label_weights) hoists indices_array and valid_mask.
The function takes them positionally and uses indices_array.size for
size checks, removing three len(indices) calls and the optional-kwarg
fallback paths.
* refactor(weights): pipeline API consolidation pass
- compute_label_weights: drop Optional placeholder, accept
Sequence[int] | NDArray[np.integer] for indices
- standardize Optional[X] -> X | None across module (PEP 604)
- _impute_weights: positional call instead of keyword on single arg
- _pivot_equivalent_count: remove unreachable threshold <= 0 branch
(survivors.size > 0 implies survivors.max() > 0 because the input
has been sanitized to non-negative values upstream)
- _scatter_weights: drop dead 'if not np.any(valid_mask)' early
return; vectorized assignment is a no-op when the mask is all-False
- sanitize_and_renormalize: clarify empty-input semantics in docstring
* fix(weights): zero leading and trailing non-finite runs in _impute_weights
The boundary mask only covered the strict tip positions (index 0 and -1),
so multi-element non-finite runs at the boundary were median-imputed
instead of zeroed. With input [NaN, NaN, 1.0, 2.0, NaN, NaN] the function
returned [0.0, 1.5, 1.0, 2.0, 1.5, 0.0] instead of [0, 0, 1.0, 2.0, 0, 0],
silently extending pivot weight to the unconfirmed boundary candles.
Use np.argmax on the finite mask to detect the leading and trailing
non-finite runs and zero the entire run, matching the docstring contract.
* fix(weights): floor stacked metrics in geometric and harmonic aggregation
Power means with p<=0 collapse to 0 on a single zero in the stack:
pmean([1, 0, 3], p=-1) = 0.0 and pmean([1, 0, 3], p=0) = 0.0. Combined
with compose_sample_weights' (arr <= 0) drop_mask predicate, a single
metric returning 0 on a pivot silently drops that row entirely.
Floor stacked_metrics at np.finfo(float).tiny only inside the
geometric_mean and harmonic_mean branches so all-positive pivots
survive aggregation. arithmetic_mean, quadratic_mean, weighted_median
and softmax branches are untouched.
* fix(weights): log when out-of-range pivot indices are dropped
compute_label_weights silently filters out pivot indices outside
[0, n_values) via valid_mask. This made upstream contract violations
invisible: a stale or off-by-one index list would simply produce zero
training weight on those rows with no diagnostic.
Emit logger.warning with the count and dropped fraction whenever
n_dropped > 0 so the upstream caller can spot the issue.
* refactor(weights): collapse 4x label-config validators into a registry
Replace the four near-identical _validate_*_params + get_label_*_config
pairs with a single _LABEL_KIND_REGISTRY mapping each kind name to
(specs, defaults). _label_kind_validator builds the validator on the
fly and get_label_kind_config dispatches to _get_label_config with the
appropriate spec/default pair. The four public get_label_*_config
helpers remain as thin wrappers so existing callers in QuickAdapterV3
and QuickAdapterRegressorV3 are unaffected.
_LabelTransformerConfig.from_dict (LabelTransformer.py) is intentionally
out of scope: it would require propagating a logger through
BaseTransform's freqtrade-side interface, which is upstream-controlled.
* docs(weights): drop verbose empty-input note from sanitize_and_renormalize
The added sentence paraphrased the existing collapse line and restated
obvious facts about zero-length vectors without contractual information.
Revert to the concise pre-W1 docstring.
* docs(weights): drop get_label_kind_config docstring for consistency
The 4 sibling get_label_*_config wrappers have no docstrings; their
parameter names and the registry name self-document the contract. Drop
the redundant docstring on get_label_kind_config to match the family
style.
* fix(weights): drop tiny floor in geometric/harmonic aggregation
The floor at np.finfo(float).tiny added in b1f86a0 preserved pivots
whose metrics included an exact zero, but a zero metric is itself a
'signal absent' marker that downstream compose_sample_weights drops
via the (arr <= 0) mask. Floor was masking the intended drop.
Restore the upstream pmean behavior so a zero in any geometric or
harmonic input produces an exact 0.0 combined weight, allowing
drop_mask to drop the pivot as designed.
Jérôme Benoit [Mon, 25 May 2026 16:05:22 +0000 (18:05 +0200)]
feat(quickadapter): log new label_weighting fill_* tunables
Mirror the existing softmax_temperature conditional-logging pattern:
always log fill_method, then log fill_epsilon and fill_epsilon_baseline
only under fill_method == "epsilon", and fill_sigma_candles only under
fill_method == "gaussian". Keeps the per-column "Weighting:" log block
consistent with the resolved config and avoids printing tunables that
have no effect under the active mode.
Jérôme Benoit [Mon, 25 May 2026 15:52:05 +0000 (17:52 +0200)]
feat(quickadapter): add soft off-pivot weighting (epsilon, gaussian) to label_weighting (#74)
* feat(quickadapter): add soft off-pivot weighting (epsilon, gaussian) to label_weighting
Adds three off-pivot weighting modes behind a new fill_method tunable in
freqai.label_weighting:
- zero (default): current hard-zero behavior, retained for backward
compatibility.
- epsilon: off-pivot rows receive a flat baseline
fill_epsilon * <baseline>(pivot_weights), where <baseline> is mean or
median, controlled by fill_epsilon_baseline.
- gaussian: off-pivot rows receive a per-row weight from a heatmap-style
decay max_p w_p * exp(-(i-p)^2 / (2 sigma^2)), controlled by
fill_sigma_candles (>= 0.5).
The default is zero so existing configs without the new keys behave
identically. Switching fill_method materially changes per-leaf weight
mass and may require GBM hyperparameter retuning; flagged in the README
description column.
Implementation:
- Adds FillMethod/FILL_METHODS and FillEpsilonBaseline/FILL_EPSILON_BASELINES
Literal types and tuples in LabelTransformer.py.
- Extends DEFAULTS_LABEL_WEIGHTING with the four new keys and their
defaults.
- Extends _WEIGHTING_SPECS in Utils.py with corresponding _EnumValidator
and _NumericValidator entries (epsilon in [0, 1], sigma_candles >= 0.5).
- Refactors _scatter_weights to accept fill_weights as a precomputed
array plus optional indices_array/valid_mask kwargs; preserves
pre-existing length-mismatch ValueError and empty-input early-return
semantics.
- Adds _gaussian_fill_weights helper with in-place pipeline keeping peak
memory at one (chunk, M) buffer; chunk-by-N keyed on
_GAUSSIAN_FILL_CHUNK_BUDGET = 50_000_000 cells (~400 MB peak); emits a
density warning when M / N > 0.1; rejects negative pivot weights.
- Adds *, logger: Logger keyword-only parameter to compute_label_weights
and updates the single call site in QuickAdapterV3.py.
- Replaces the raw nonzero count in compose_sample_weights with a
pivot-equivalent count helper (_pivot_equivalent_count) so the sparse
training mass warning stays meaningful under epsilon / gaussian.
Documentation:
- Four new rows added to the README configuration tunables table under
Label weighting; fill_method flagged as requiring trained-model
deletion when changed.
- Four new keys added to config-template.json under label_weighting.
Verified manually on host via AST extraction harness (no automated test
infrastructure exists in quickadapter/):
- STATIC_OK: defaults + tuples assertions pass.
- SPOTCHECK_4..9: cluster amplification (out[50] ~= 8.0), sigma < 0.5
rejected, negative pivot weights rejected, density warning emitted at
M/N=0.2, empty pivots return zeros, mean/median epsilon ratio = 20.8x.
- SPARSE_4A..C: sparse-mass warning fires under zero mode + sparse
pivots and gaussian sigma=0.5 underflow regime; silent under broad
gaussian fills.
* fix(quickadapter): harden sanitize_and_renormalize against rescale overflow and drop_mask contract violations
Two production-quality safeguards on the load-bearing primitive used by
the new fill_method dispatch in PR #74, plus one cosmetic comment
cleanup.
1. Subnormal-total rescale overflow guard:
When the sum of sanitized weights falls into a subnormal range
(e.g. a single 1e-310 survivor among zeros, n=1000), n/total
overflows to +Inf and safe * Inf propagates Inf to every nonzero
entry, producing mean(out) = NaN and silently violating the
documented mean=1 invariant. The fix computes the rescale factor
into a local c, checks np.isfinite(c), and falls through to the
existing uniform-fallback path with a distinct warning message
('rescale factor non-finite') so operators can distinguish this
regime from the existing 'weights collapsed' case. Bit-identical
on all common paths; c -> 0 underflow is unreachable
(min c = 1/DBL_MAX > 0).
2. drop_mask shape and dtype assertions:
sanitize_and_renormalize is now load-bearing for compose_sample_weights
under all three fill_method modes (zero/epsilon/gaussian). Numpy
broadcasts a (k, n)-shaped mask silently, breaking the (n,) output
contract. Shape and dtype precondition checks raise ValueError
early with prefixed messages matching the function's existing
logger style. Dtype check uses np.issubdtype(..., np.bool_) so
any boolean alias (bool, np.bool_, 'bool') is accepted; integer
masks are rejected.
3. LabelTransformer.py: replace 'current behavior' comment with
'default' on FILL_METHODS[0] since the comparison no longer makes
sense once the PR is merged.
Verified manually:
- REVIEW_FIX_1A..C: bool, np.bool_, 'bool' all accepted; int rejected.
- REVIEW_FIX_2A..B: subnormal-overflow path emits the new distinct
warning; real collapse path emits the original warning.
- REVIEW_FIX_3_OK: docstring contradiction removed.
- REGRESSION_OK: bit-identical common path.
- All original PR #74 verifications still pass (SPARSE_4A..C,
SPOTCHECK_4..9).
* fix(quickadapter): short-circuit compute_label_weights on empty pivot weights
When metrics[strategy] is empty but indices is non-empty, the new
fill_method dispatch in epsilon/gaussian arms slices weights[valid_mask]
before _scatter_weights can short-circuit, raising IndexError on a
size-0 / N-mask shape mismatch. Pre-PR _scatter_weights returned the
default-filled array silently in this case (preserved invariant noted
inline at the empty-input early return).
Add a short-circuit before the dispatch so the contract is consistent
across all three fill methods.
Also trim _gaussian_fill_weights docstring to match the codebase style
(neighboring private helpers carry no docstring or a single short
paragraph) and drop a redundant in-line comment that the in-place
np.multiply(out=buf) pattern already conveys.
Verified on the AST-extraction harness (pre-fix reproduction → fix
verification): 12 contract assertions across 4 edge cases x 3 fill
methods, plus crossmode + non-empty differentiation, all pass; PR #74
SPOTCHECK_4..9, SPARSE_4A..C, REVIEW_FIX_*, REGRESSION_OK still pass.
* chore(quickadapter): bump strategy and regressor version 3.11.10 -> 3.11.11
* refactor(quickadapter): polish label_weighting docs, comments, and sparse-mass diagnostic
Three coordinated polish edits following final-review feedback:
1. _pivot_equivalent_count: replace the 0.5 * median threshold with
_PIVOT_EQUIVALENT_MAX_FRACTION * surviving max (default 0.1). The
median-based heuristic saturated at N under epsilon mode (off-pivot
floor dominates the median once N >> M), silencing the warning the
docstring claimed to provide. The max-relative threshold separates
pivot-class rows from off-pivot fill across the bimodal regimes
fill_method introduces. Constant is module-level and named so the
choice is auditable; warning text now self-describes the threshold
('rows above 10% of surviving max').
2. _scatter_weights: trim the 'Order matters...' comment from 3 lines
to 1 line. The shorter form pins the intentional ordering without
paraphrasing git history; future 'validate inputs first' refactors
are still flagged.
3. README: extend the fill_method row with a concise retuning hint
(per-leaf regularization + Optuna study reset) so the operator
guidance surfaces in user docs, not only in the planning artefact.
Tighten fill_sigma_candles description to match neighboring-row
density.
Verified manually:
- SPARSE_4A..C: original PR cases still pass.
- SPARSE_4D: epsilon+sparse pivots (M=20, N=1000) now correctly fires
the sparse-mass warning (was silenced with median-based threshold).
- SPARSE_4E: zero+skewed pivots ([1,1,...,10]) still fire under the
new threshold (no regression on the skew case).
- SPOTCHECK_4..9, BUG_74_FIX_*, REVIEW_FIX_*, REGRESSION_OK: all
unchanged.
Jérôme Benoit [Mon, 25 May 2026 12:01:55 +0000 (14:01 +0200)]
style(weights): use !r repr for strategy in compute_label_weights ValueError
Replace hardcoded 'none' literal with f"{strategy!r}" to match the
codebase convention of !r-formatting external string identifiers in
error messages and avoid silent drift if WEIGHT_STRATEGIES[0] changes.
Jérôme Benoit [Mon, 25 May 2026 03:18:52 +0000 (05:18 +0200)]
fix(plot): switch raw direction/weight to bar with steelblue color
Raw discrete signals (direction +/-1/0, weight at pivots) are sparse and
read better as bars than lines. Steelblue contrasts with the orange
smoothed lines for clearer visual separation.
Jérôme Benoit [Mon, 25 May 2026 02:38:25 +0000 (04:38 +0200)]
fix(weights): canonical sanitize_and_renormalize and compose_sample_weights
Derived from independent dual-oracle mathematical specification with
proofs (mean=1 invariant, drop preservation, idempotency, collapse
degradation chain).
sanitize_and_renormalize:
- Fix latent bug: fallback path with non-empty drop_mask returned ones
zeroed at drop_mask but did not renormalize, breaking the mean=1
contract. The fallback now renormalizes so mean(out) == 1 holds on
surviving rows.
- Replace .copy()+mutation with np.where for drop_mask application.
compose_sample_weights:
- Replace the post-compose combined.sum() guard (which duplicated the
predicate sanitize_and_renormalize re-evaluates internally) with a
single survivor-aware predicate covering drop_mask | ~isfinite | <=0
in one pass. The check is the explicit branch point for the base-
weights fallback when the label-weighted product collapses on
surviving rows; this preserves the recency signal and the label-
derived drop_mask instead of degrading to uniform.
- Warn when nonzero/n falls below SPARSE_TRAINING_MASS_THRESHOLD (5%,
module-level constant) so operators can spot the sparse-training
regime that pivot-only weights produce on long series with few pivots.
QuickAdapterV3._log_strategy_configuration:
- Warn at startup when label_smoothing.method is 'smm' or 'savgol'
(with polyorder>=2) combined with a non-'none' label_weighting
strategy, since these kernels can collapse a sparse weight signal
and trip the all-rows-dropped guard.
Jérôme Benoit [Mon, 25 May 2026 01:39:28 +0000 (03:39 +0200)]
fix(weights): pivot-only sample weights when label_weighting is active
Replace the full-series median fill with 0.0 in compute_label_weights so
non-pivot rows carry no sample weight when a label_weighting strategy is
configured. The median fill predates PR #72: when weights multiplied the
label, the fill was inert (label=0 × median=0). Once weights became the
sample_weight kwarg of model.fit, the fill silently leaked the median
into the training loss for every non-pivot row, diluting the pivot
detection signal the model is being trained for.
Concretely, training now concentrates on pivots and their smoothed
neighborhoods (via label_smoothing), and the raw/smoothed weight plots
both render clean profiles starting from zero.
strategy='none' (the default) is unaffected: compute_label_weights still
returns a uniform 1.0 vector and every row contributes equally.
Jérôme Benoit [Mon, 25 May 2026 00:45:18 +0000 (02:45 +0200)]
fix(weights): guard DI_values None and read label_frequency_candles from freqai_info
- fit_live_predictions: pred_df.get('DI_values') returns None when
feature_parameters.DI_threshold is 0 or absent (the default), causing
AttributeError on the subsequent .mean()/.std() calls. Fall back to
zeros instead.
- _label_frequency_candles: read from self.freqai_info['feature_parameters']
(matching every other access in the file) instead of self.config, which
is the top-level config dict and never contains feature_parameters
directly. The previous code silently ignored user-provided values and
always fell back to the default 'max(2, 2 * len(self.pairs))'.
Jérôme Benoit [Mon, 25 May 2026 00:31:49 +0000 (02:31 +0200)]
fix(weights): tighten observability and edge-case handling in label pipeline
- sanitize_and_renormalize accepts logger/context kwargs and warns on
uniform-fallback collapse; six call sites in QuickAdapterRegressorV3
thread their stage label (train_test_split / post_feature_pipeline /
timeseries_split, train|test).
- Warn at startup when label_prediction.method='none' for any label, since
populate_entry_trend would silently never trigger.
- Replace .notna() with np.isfinite() in the smoothed-weight clip so +Inf
produced by smoothing kernels is also zeroed instead of relying on the
downstream drop_mask in compose_sample_weights.
- _impute_weights tracks boundary NaN separately so injected zeros do not
bias the interior median; finite endpoints are now preserved.
Jérôme Benoit [Mon, 25 May 2026 00:04:41 +0000 (02:04 +0200)]
refactor(weights): collapse compose_sample_weights to single-target API
LABEL_COLUMNS is single-target by design, so the dict-shaped per-label
map and row-wise aggregation in compose_sample_weights were dead
plumbing. Flatten the signature to a single label_weights vector and
read LABEL_COLUMNS[0] directly in _compose_per_row_weights. Drop the
duplicate-column guard (unreachable under single-target). Align caller
naming on base_weights to match the callee parameter. Add a defensive
check that LABEL_COLUMNS[0] is in dk.label_list to fail loudly if the
project label constant ever diverges from freqtrade's runtime view.
Jérôme Benoit [Sun, 24 May 2026 23:44:44 +0000 (01:44 +0200)]
refactor(quickadapter): drop unused sample_weighting tunables
The sample_weighting.{aggregation,softmax_temperature} options were inert:
LABEL_COLUMNS is single-label by design and compose_sample_weights only
ever sees one per-label vector, making row-wise aggregation an identity
operation regardless of the configured mode.
Removes the config block, validation specs, getter, and README entry;
compose_sample_weights keeps its kwargs with safe defaults (arithmetic_mean,
T=1.0) so the call site stays trivial.