ANULUM / RESEARCH

KYMA

Trainable abstractions for checkable AI reasoning and planning. A proposed research programme with explicit evidence, tests and limits.

Preparing an application · Updated · Proposed duration: 36 months

Public scientific-page draft v1 · 2 October 2026

Miroslav Šotek · Marbach SG, Switzerland · ORCID 0009-0009-3560-0851 · protoscience@anulum.li

Status: preparing an application to EIC Pathfinder Challenges 2026 — DeepRAP. The application has not been submitted and no funding has been awarded. The 36-month duration, research objectives, personnel effort, work packages, collaboration offers and output quantities below are proposed. Quantitative success criteria are future targets. Existing preliminary findings are identified separately and retain their task-specific limitations. The later dynamics ablation is reported qualitatively until its exact-revision public replication bundle is verified.

Official call and guidance

Project overview

KYMA asks whether abstractions learned as phase relationships can transfer to new compositions of tasks through typed interfaces and independently checked reasoning or planning. Its proposed system combines trainable phase representations, specialised language-model interfaces, conventional controls and explicit acceptance/refusal checks. The programme separates a possible representational benefit from any benefit of integrated oscillator dynamics. It also measures the resources needed for correct, admitted task completion.

The proposal builds on research software in the Anulum ecosystem. Relevant starting components include SCPN-QUANTUM-CONTROL for phase dynamics and preliminary probes, SCPN-PHASE-ORCHESTRATOR for signal/phase analysis, SC-NeuroCore for bounded numerical/RTL work, SCPN-CONTROL for checked control interfaces, Director-AI for grounding diagnostics and Synapse for research coordination. These are starting capabilities, not evidence that the proposed integrated KYMA system already exists.

Scientific programme

1.1 Objectives and relevance to the Challenge

The problem KYMA addresses. A cognitive system must compose known relations in unfamiliar situations, justify its outputs and account for the resources used. Strong neural and neuro-symbolic methods already solve parts of this problem; KYMA tests whether a learned phase representation adds transferable abstraction beyond them. Its defining question is: can an abstraction learned in one task become a reusable, independently checkable building block for another, within a measured resource budget? Abstraction is the primary scientific focus; reasoning and planning test its usefulness in the two pilots.

The central idea. KYMA (from the Greek word for "wave") tests a different substrate: a network of coupled oscillators whose couplings are trained by gradient descent through a differentiable integrator. In such a network a relation between two groups of units is a phase-locking pattern (in-phase, anti-phase, or a fixed phase lag), a concept is a re-usable synchronisation motif, and a proposed conclusion is a read-out of the learned phase state. Integrated dynamics and an analytic phase map are separate candidate implementations. The design offers an explicit state for predicate read-out, a bounded interface for formal analysis and a measurable implementation cost. These are design opportunities, not established advantages over other representations. The learned motifs, read-out stability and benefit of dynamics are tested separately; digital simulation does not imply an energy advantage or an available physical implementation.

Figure1 — integrated architecture with matched conventional comparison.
Figure1 — integrated architecture with matched conventional comparison. Proposed architecture; matched conventional controls use the same checking and resource boundaries.
Open the full-size figure

Architecture: language/document/numerical-feature input → specialised typed interface → learned oscillator motifs and symbolic constraints/plans → independent checker → evidence-bounded result; fixed conventional comparison uses the same sources/tools/checker.

Preliminary evidence. Before writing this proposal we ran a test whose protocol was committed before the run of the core assumption (Section 1.3.2): on held-out combinations of two learned relations, a 17-oscillator substrate with per-relation coupling gating reached 80.1 % ± 3.0 % accuracy (5 seeds), against 36.9 % for a parameter-matched MLP, 50.7 % for a code-conditioned graph neural network and 24.2 % chance. The MLP fits its training set perfectly (100 %) and still fails on the new combinations; an ablation without gating collapses to 33 %. A first version of the test failed, and we publish that negative result and its diagnosis alongside. That result is toy-scale and in-class. A subsequently completed symbolic-ground-truth probe reached 100.0 % versus 44.1 % for the best matched baseline on one held-out ordered composition, five seeds and 64 shared test cases (Section 1.3.2). It removes the oscillator-teacher dependence for that task, but a completed frozen ablation found the dynamics-free fixed-point map equally accurate. Under its predeclared attribution rule, the v3 result is carried by the staged phase-arithmetic structure, not by oscillator dynamics. KYMA tests transfer beyond this encoding and any separate benefit of integration.

Overall objective. Establish whether trainable oscillator abstractions, specialised language-model interfaces and independently checked reasoning/planning can form one reproducible cognitive system that transfers to unseen compositions and task conditions under bounded training and deployment resources, demonstrated at TRL4 in two planned pilot domains.

The central scientific test is learned abstraction: SO1 tests reuse on unseen compositions, and SO6 tests its value in checked pilot workflows. SO2/SO3 establish the interface and its acceptance boundary; SO4 tests planning; SO5 measures integrated deployment. Frontier comparisons, FPGA kernels and dissemination support these tests. Scientific success and valid work-package completion remain separate: a completed negative report does not meet a superiority target.

# Specific objective (capability) Acceptance criterion (protocols committed before the run in D2.1 / D3.1 / D4.1 / D5.1) WP / month
SO1 Compositional abstraction on the substrate (Deep Abstraction) — scale the learned phase representation to ≥ 128 phase units (oscillators in the integrated variant) and ≥ 8 relation types; train on tasks whose ground truth is not generated by an oscillator system: (a) a relation algebra with held-out compositions, (b) the NoRA relational-reasoning benchmark, (c) a ConceptARC subset re-encoded as relational scenes On held-out compositions the phase representation exceeds the strongest validation-selected conventional learned comparator (including adequately trained structured controls, with matched data and search budgets) by ≥ 10 percentage points in ≥ 2 of the 3 families, five seeds, paired source/composition-group analysis with multiplicity controlled across the three families; protocol frozen before training; NoRA results reported next to the published numbers for neuro-symbolic, transformer and GNN models. MS1 go/adapt gate at M12. WP2, M12/M18
SO2 Motif–symbol interface with certificates (Deep Reasoning) — read predicates from synchronisation states; compile symbolic constraints back into coupling structure; perform deductive and causal inference over read-outs Predicate read-out ≥ 95 % on in-distribution states; all outputs labelled accepted carry an independently valid certificate; the target is certificate yield ≥ 80 % on the frozen reference-solvable inference set, with semantic correctness reported separately; on CLadder (causal), ZebraLogic-class constraint puzzles and FOLIO (first-order logic), accuracy, certificate rate and J/reference-correct completion reported against a Scallop-class neuro-symbolic baseline at matched compute (accuracy targets frozen in D3.1; published best results — 95.3 % CLadder with fine-tuned LLMs — stated as context, not as the bar). WP3, M30
SO3 Verified trust spine (trustworthiness, formal guarantees) — read-out, acceptance and refusal components specified and model-checked; reasoning traces hash-chained; robustness measured by an independent team 100 % of properties in the D4.1 catalogue resolved by proof or counterexample; every deployment-critical acceptance/refusal property must be proven for the admitted configuration, otherwise that configuration remains disabled; the implemented acceptance path cannot label an output accepted without a certificate passing the specified checker (proof under the D4.1 model assumptions; factual/document support measured separately); on an AgentDyn/AgentDojo-class attack suite, targeted attack success ≤ 2 % with utility ≥ 90 % of the undefended system (security and over-defence reported together); external red-team report published. WP4, M34
SO4 Planning as control of slow variables (Deep Planning) — hierarchical, contingency and continual re-planning realised as control over the substrate's slow variables in domains whose transition model is learned from trajectories, not given; every plan exported to PDDL and checked by a validator against the true domain 100 % of executed plans validated (invalid plans are refused, never executed); on Blocksworld and Mystery Blocksworld with the model learned from ≤ 200 example trajectories, target ≥ 10 pp valid-task-coverage gain over the strongest validation-selected learned-model comparator at equal evaluation budget; include a classical planner supplied the same learned model. Compare contemporary frozen LLM/tool baselines; historical o1-preview timings are context only. The true-domain planner is a privileged reference, never a training oracle. A sealed evaluator checks the submitted candidate against the true model; it does not return oracle validation feedback for candidate search. Planning and permitted replanning receive only the common observation/action budget. Freeze the number of proposals and observations before testing; count invalid/refused proposals in coverage. Re-planning bounds are fixed per task before evaluation. WP5, M33
SO5 Integrated TRL-4 system, energy protocol and shared benchmark (integration, "constrained computational resources") SO1–SO4 integrated on one metered node (≤ 1.5 kW), with a bounded SC-NeuroCore FPGA comparison; two pilots run end-to-end (Section 1.3.3); every headline comparison reports cost and its resource boundary; metered local runs report J/reference-correct completed task, while unavailable cloud energy is marked unknown; ≥ 3 task families released as a FAIR contribution to the DeepRAP benchmark. MS5 TRL-4 validation at M36. WP6, WP7, M36
SO6 Specialised language-model training and transfer — two specialisations connect language to learned motifs, scientific tools and checked plans Matched conventional checked specialisation versus KYMA hybrid; five seeds and held-out composition groups. Target point-estimate gain ≥ 10 pp in each pilot and multiplicity-adjusted confidence interval excluding zero; paired grouped analysis, five-seed variability and Holm-adjusted significance across two endpoints (family-wise 0.05). Interval construction is frozen in D3.1; no pooling of the two domains to hide a failure. Report failures, refusal/utility and energy separately. Model packages are due at M24; interface and transfer reports follow at M30/M33. The confirmatory two-pilot endpoint is evaluated in WP6 at M34 on an untouched cohort. Artefact/report completion does not require a positive hypothesis. WP3/WP5/WP6, M30/M33/M34

Graceful degradation (pre-committed). If SO1 fails the MS1 gate, oscillator layers are repositioned as abstraction modules inside conventional pipelines and SO2–SO4 continue on that basis; SO3 (verified trust spine) and the energy protocol of SO5 are useful to the portfolio regardless of the substrate result. A negative substrate result remains publishable, but does not establish the central breakthrough. The fallback retains the contracted reasoning, planning and integration evaluations; it reports precisely which objectives were met and which were not.

Relevance to the Challenge objectives and expected outcomes.

DeepRAP element (WP 2026 / Challenge Guide) KYMA contribution
"novel frameworks … inspired by … physics" (specific objectives) Oscillator synchronisation as the learning and reasoning substrate (SO1)
Deep Abstraction: generalise from limited data, concept formation, world models Re-usable motifs; held-out-composition tests with sample counts reported (SO1)
Deep Reasoning: causal and logical inference, explainable rationales; neuro-symbolic encouraged Motif–symbol interface; certificates checked independently (SO2)
Deep Planning: hierarchical, contingency, continual re-planning, formal guarantees Control of slow variables; validator-checked plans (SO4)
Expected outcome: multimodal data and knowledge, uncertainty Time series (Pilot A trajectories), relational scenes and graphs (SO1), text documents (Pilot B); phase coherence (order parameter) is used as a confidence signal for refusal, and its calibration is measured, not assumed (SO2, SO3)
Expected outcome: constrained computational resources Single metered inference node; J/correct task within the measured boundary, cost and separately scoped cloud/training resources (SO5)
Expected outcome: provable trustworthiness, EU AI Act Model-checked trust spine, independent red-team (SO3); AI Act mapping (D1.3)
Expected outcome: TRL-4 cognitive system on real tasks Two pilots on one integrated system (SO5)
Expected outcome: new methods and metrics for evaluating reasoning, trustworthiness and compute use Certificate rate, reference-correct valid plans/joule, composition gap (SO2, SO4, SO5)
FAIR; synergies with TEFs, EBRAINS, RAISE, AIoD, Quantum Flagship Section 1.3.6 (FAIR) and 2.3 (named synergies)

Specific conditions. The Challenge admits single applicants including natural persons; KYMA is prepared for submission by one natural person established in Switzerland (associated country; Swiss entities participate as beneficiaries since the 2025 programme year). Portfolio activities are planned in a dedicated work package of 17 person-months (WP7, following Annex 1 of the Challenge Guide).

Portfolio categories (self-assessment, Challenge Guide §3.1).

Category Focus Secondary focus Evidence / justification
Cognitive capability Deep Abstraction (DA) Multi-Capability (MC) Deep Abstraction is the primary scientific contribution; reasoning and planning evaluate the learned representation in an integrated system
Technological approach Novel Interdisciplinary Frameworks (NIF); Neuro-symbolic AI (NeSy) Formal Methods Integration (FMI); Trustworthiness Mechanisms Integration (TM) Physics-inspired substrate with symbolic read-out; model-checked trust spine; independent red-team
Application domain Scientific Discovery (SD) Decision Support Systems (DSS) Pilot A: structure discovery in coupled dynamical systems; Pilot B: grounded decision support with refusal
Synergies Benchmark Development (Ben); Interoperability (Int) Joint Pilots and Demonstrations (JPD) Benchmark families and energy protocol (WP7); open read-out/certificate interface; offer of the substrate as a module in joint pilots

Portfolio reuse — a versioned typed-task/result schema, numerical witnesses, independent checker, composition benchmark and resource-boundary record let another project replay a checked task without KYMA weights. WP7 supplies a runnable example and per-asset rights manifest. Raw results and frozen protocols support comparisons. Benchmark co-leadership and trust-method contributions are proposed; neither appointment nor uptake is assumed.

1.2.1 State of the art and its limits

Benchmark evidence motivates controlled comparisons. ARC-AGI-2 tests passive abstraction, whereas ARC-AGI-3 introduces interactive tasks 1; model revision, harness, task split and per-task expenditure must accompany comparisons. Their leaderboards report monetary cost rather than metered energy or a certificate of correctness. Planning and travel benchmarks illustrate the importance of obfuscation, explicit constraints and solver-assisted checking [2–4]; ZebraLogic examines increasing logical complexity 5. These are distinct task populations, not a common ranking of general reasoning ability. KYMA therefore compares representations on shared frozen tasks, with the same information, checker and resource accounting, rather than extrapolating from headline leaderboard percentages.

Neuro-symbolic AI supplies structure but keeps a statistical substrate. Probabilistic-logic programming and its GPU-compiled successors (DeepProbLog, Scallop, Dolphin, Lobster, DeepProofLog) make symbolic inference differentiable and fast [6–8]; neuro-symbolic concept composition addresses vision-language generalisation on ReaSCAN 9; meta-learning can make transformers systematic 10. These approaches show why conventional structured representations and symbolic solvers are strong comparators. KYMA investigates learned phase motifs as an alternative representation; it does not assume that existing neuro-symbolic methods lack meaningful internal structure. Recent benchmark work warns that path-composition tasks are too easy for such architectures and proposes harder relational reasoning (NoRA) 11.

Oscillator computing has demonstrated structured tasks. Synchrony-based models bind features into objects (Complex AutoEncoder, Rotating Features, KomplexNet) 12; oscillatory recurrent and state-space models learn long sequences stably (coRNN, LinOSS, HORN) 13; equilibrium propagation trains physical oscillator networks robustly to frequency dispersion, so far on MNIST-class classification 14; oscillator arrays solve combinatorial optimisation such as Max-3-SAT 15. Three recent works go further. AKOrN replaces threshold units by Kuramoto oscillators and solves Sudoku (100 % in-distribution, 89.5 % on harder unseen boards with extended test-time integration and energy-based voting) 16; KoPE adds a Kuramoto phase state to vision transformers and reports benefits on few-shot ARC 17; Phasor Agents use oscillator graphs with replay for maze navigation 18. KYMA evaluates a specific combination: held-out relation composition, typed predicate read-out, an independently checked acceptance boundary and metered task performance. The claimed advance is tested against oscillator and conventional structured comparators; these cited works are related prior art, not evidence that every component is individually unprecedented. EU-funded oscillator research (THERMODON, PHASTRAC, RadioSpin, NeurONN) targets hardware and perception-level learning; neuro-symbolic projects (DeepLog, TUPLES) use statistical substrates; the cited CORDIS examples establish related hardware and neuro-symbolic efforts, not an exhaustive absence of competing research 19.

Verification requires an explicit implementation boundary. The cited VNN-COMP/ONNX verification ecosystem 20 motivates a separate specification of KYMA's numerical dynamics and implemented read-out; absence from a particular benchmark is not evidence that dynamical-system verification is unexplored. Provable agent security is emerging at system level (CaMeL: 77 % of AgentDojo tasks with provable security vs 84 % undefended) 21, while AgentDyn reports security/over-defence trade-offs among ten evaluated defences in dynamic tasks 22.

Task-level resource measurement builds on existing work. ARC reports cost in dollars; MLPerf and AI Energy Score measure inference energy 23. Panigrahy and Tyagi already define energy per successful goal across agentic workflows, including retries and failures 24; Intelligence per Watt also connects quality and resource use 25. KYMA does not claim invention of energy per solved task. Its contribution is a reproducible common-node protocol for reference-correct, checker-admitted completion, explicit measurement boundaries and comparisons of phase-hybrid and conventional implementations, offered for portfolio interoperability.

1.2.2 The breakthrough: a substrate that is at once a learning medium and a symbol medium

KYMA's proposed breakthrough is a learned phase representation that can be reused through a checkable symbolic interface. A continuous state is not itself a discrete symbol: the specified read-out produces predicates, whose stability and composition must be established experimentally. Specialised LLM training learns the interface to this representation; it is not a separate foundation-model programme. The research claim concerns the learned phase abstraction and its transfer through a typed interface under controlled comparisons. Integrators, low-rank adaptation, tool use, validators and energy accounting are established techniques; their presence alone is not a breakthrough. The five components and cross-cutting measurement contribution below separate the proposed advance from its enabling methods:

Component Nearest prior art What is new in KYMA
C1 Compositional motifs (SO1) AKOrN (Sudoku), synchrony binding (objects) Per-relation gated couplings that compose on unseen relation combinations, tested against matched MLP/transformer/GNN baselines on internal and standard relational suites (NoRA)
C2 Motif–symbol interface (SO2) Neuro-symbolic programs over neural embeddings Predicates read from, and constraints compiled into, the learned phase state; numerical witnesses replayed separately from independent encoded-constraint/proof checks
C3 Verified trust spine (SO3) Feed-forward verifiers; system-level agent security Scoped verification of finite-precision read-out, acceptance and refusal; stated assumptions bind proofs to admitted implementations
C4 Planning as control of slow variables (SO4) Classical/learned-model planners, checked LLM/tool planners and oscillator maze agents Test whether transferred phase motifs improve learned-model planning relative to matched conventional transition representations/search; slow-variable control is compared separately, and every proposed plan is validated
C5 Specialised language–abstraction learning (SO6) Conventional post-training and typed tool use Controlled learning of interfaces to oscillator abstractions, with matched conventional representations, held-out transfer and independent output checking; gain remains a hypothesis
P1 Metered checked-task performance (SO5) Goal-level energy accounting 24; quality/resource profiling 25 Common-node energy per reference-correct completed task, all attempts included; reproducible protocols and explicit cloud boundaries for portfolio use

Discriminating evidence. Shared encoding, sources, checker and validation-only model selection compare phase representations with adequately trained conventional structured controls. An analytic phase map isolates the phase prior; recurrent/filter controls receive the same observation history, memory and budget. Relation renaming, unseen compositions and size/noise changes are frozen before testing. NoRA 11 tests beyond phase-friendly path algebra. D2.1 fixes a separate noisy/time-varying dynamics endpoint, tolerance, budget and analysis before its runs; clean-task success cannot substitute for it.

Result under the frozen comparisons Permitted scientific conclusion
SO1 transfer criterion met; dynamics endpoint met Evidence for representation transfer and an additional dynamics benefit, within the tested scope
SO1 met; analytic phase matches integration Representation transfer supported; dynamics benefit unestablished
Dynamics endpoint met; SO1 not met Scoped robustness benefit only; no broad abstraction breakthrough
Neither met, or precision insufficient Negative or inconclusive result; activate the recorded adaptation without relabelling it as success

If a phase-free structured control reaches ceiling, phase quality superiority is not established. Resource, sample-efficiency and robustness endpoints remain distinct; favourable secondary results cannot rescue a failed primary claim.

Ambition, stated with its limits. We do not aim to beat frontier models on ARC accuracy; the proposed TRL 1–4 research is evaluated on transferable abstraction and resource-bounded checked task completion. We aim at the axes the Challenge names and that current systems leave open: generalisation to new combinations from little data, verifiable conclusions, guaranteed plans, and measured resource use. The preliminary probes (Section 1.3.2) provide toy-scale evidence under explicit architectural priors; KYMA is the programme that finds out whether it holds at the scale of a TRL-4 system.

1.3 Plausibility of methodology

1.3.1 Concepts, models and assumptions

Substrate. Units are phase oscillators with natural frequencies ω_i and state θ_i. The dynamics are Kuramoto-class, dθ_i/dt = ω_i + Σ_j K_ij(c) sin(θ_j − θ_i − α_ij), integrated by a fixed-step Runge–Kutta scheme that is differentiable end-to-end. The symbolic v3 probe additionally uses staged writes, reset and triadic phase-sum terms; it is not evidence that the pairwise equation alone solves that task. WP2 compares pairwise, triadic, staged dynamics and an analytic fixed-point map, with the same encoding and trainable gates. The coupling K(c) = K_base + Σ_r c_r ΔK_r is gated by a relation code c: each relation r owns a coupling increment restricted to the unit groups it relates. This gating was the decisive change between our failed and our successful probe (Section 1.3.2).

Read-out. Relations are read from order parameters and pairwise phase differences after integration or the specified analytic phase map; predicates are thresholded functions of these quantities. SO3 verifies specified properties of the implemented read-out under bounded numerical/noise assumptions; proof of the continuous dynamics or of a natural-language interpretation does not follow automatically.

Abstraction hypothesis (SO1). A motif trained for one relation re-uses the same coupling increment wherever the relation applies; composition combines gated increments, and generalisation depends on the representation and specified computation of the resulting phase configuration. The v2.1 ablations support initial-state-dependent phase computation within that in-class task. The separate symbolic v3 result establishes learned composition under a phase-lattice/staging prior. Its completed ablation isolates the fixed-point phase map: it matches integrated accuracy, so this result does not support a dynamical mechanism. Two hypotheses are separated. Hrepr asks whether learned phase structure transfers beyond its encoding better than strong conventional learned representations. Hdynamics asks whether integration adds robustness on noisy, partially observed or time-varying tasks beyond the same analytic phasor and suitable recurrent/filter controls, matched for observation history, memory and budget. Analytic phasor and structured program controls expose the built-in algebraic prior. If a permitted algorithmic control reaches ceiling, quality superiority is not claimed; predeclared secondary resource/data/robustness endpoints remain separate. A representation gain without a dynamics gain is reported as such; failure of both is a comparative negative result, not the proposed breakthrough.

Reasoning hypothesis (SO2). Candidate encodings map symbolic constraints to coupling structure, for example exclusion to anti-phase frustration or implication to a directed phase lag. These are proposed constructions: phase locking alone does not establish logical implication or constraint satisfaction. WP3 tests encoding/read-out soundness against independent symbolic references, retains counterexamples and admits only the properties established for the chosen encoding. Relaxation proposes a state; independent constraint/proof checks decide the declared logical claims. Active constraints and numerical witnesses document the proposal without replacing those checks.

Planning hypothesis (SO4). Fast variables (phases) settle into motifs; slow variables (coupling gains, natural frequencies, gating codes) select which motif sequence unfolds. Planning is posed as optimal control of slow variables; contingency plans are pre-shaped basins; re-planning is receding-horizon control after a perturbation. Plans are emitted in PDDL and checked by an independent plan validator, so plan validity against the supplied model does not depend on trusting the proposal generator. Correctness of the model and its real-world applicability remain separate checks.

Key assumptions and how each is tested early. A1: gated motifs keep composing when relation count and network size grow (tested by M12, MS1). A2: tasks whose ground truth is not oscillatory can be embedded without losing the advantage (M12, MS1 — the most important risk). A3: predicates can be read out stably under noise (M18). A4: slow-variable control is trainable within the energy envelope (M24). Each assumption has a falsification criterion committed before the run; failures are published.

1.3.2 Preliminary results — task-specific findings and limitations

Probe Design Result (five seeds unless stated) Status
v1 (2026-07-18) 16 oscillators, one shared coupling matrix, relation encoded as frequency offset Substrate failed on held-out conjunction; the MLP succeeded (owner-reported early negative probe; exact percentages withheld) Negative. Diagnosis: a single shared coupling cannot realise in-phase and anti-phase on the same pair; the additive code made the task linearly decodable
v2 (2026-07-21) 17 oscillators, per-relation gated coupling, passive read-out unit, readout that depends on the initial phases Substrate 80.1 % ± 3.0 %; parameter-matched MLP 36.9 % ± 1.5 %; chance 24.2 % Pass of the frozen contract (≥ 60 % and ≥ 20 pp over MLP)
v2.1 (2026-07-21) Ablations and stronger baselines, leave-one-out over 6 held-out choices No gating 33.1 %; deep MLP 41.1 %; 4× MLP 38.7 %; GNN 50.7 %; MLP train accuracy 100 %; LOO 6/6 splits, three seeds per split, +43.3 pp 4 of 5 predictions met; one refuted → mechanism statement corrected
v3 (freeze 2026-09-29; completed 2026-09-30) Symbolic Z4 program labels; 2,112 training cases; one held-out ordered pair/query, 64 shared test cases; staged substrate 108 parameters, matched MLP/GNN/transformer Substrate 100.0 % ± 0.0 %; GNN 44.1 % ± 9.9 %; MLP 25.6 % ± 4.1 %; transformer 25.3 % ± 1.5 %; chance 25 %; five seeds, population SD Pass of frozen contract: ≥ 10 pp over best matched baseline and mean minus SD above chance. Completed ablation attributes this v3 result to the phase-arithmetic structure, not dynamics

All v1–v3 protocols were frozen before training. V1/v2 used oscillator-derived labels; v3 uses symbolic Z4 labels and a deliberately realizable phase-lattice/staging prior. No frozen run was retuned on held-out accuracy. Five seeds share 64 v3 cases, not 320 independent cases. Larger MLP/GNN diagnostics fit training perfectly but reached 23.4%/30.6% held-out; small matched baselines partly underfit.

Earlier code/data/protocols: scpn-quantum-control v1.2.0, src/scpn_quantum_control/benchmarks/kyma and benchmarks/kyma_v2 and docs/campaigns/kyma_*, Zenodo 10.5281/zenodo.18821929. V3 is separately versioned: source dbcb0b0778221f85723ab7f60118c27e6369260f; protocol freeze ddae107dc; published result commit 65808bdf2eca8e84e6d2d9ed065f8ceb3f808ad3. Its protocol is docs/campaigns/kyma_v3_symbolic_composition_prereg_2026-09-29.md; result data/kyma_v3_symbolic_composition/kyma_v3_symbolic_composition.json has SHA-256 dd27ae8953d660626abbd23f23a7ea9c5c13123fa93280bef092982104f90cb8. It is not claimed to be in the older archive.

The current proposal reports a completed, protocol-frozen dynamics ablation: the dynamics-free fixed-point phase map matched the integrated model on the same cases and seeds. The attribution rule therefore assigns the v3 result to staged phase arithmetic, including the triadic phase sum; integration and long settling were unnecessary for that task. This is an owner-reported later finding. An accessible exact-revision replication bundle has not yet been verified, so exact numerical results and timing claims from that ablation are withheld here. It is not presented as a publicly reproduced dynamics advantage.

Timing-derived CPU proxies are 4.6 J/task for the earlier substrate versus 1.3 mJ/task for MLP (nominal 15 W), and 1.68 J/item for v3 versus 0.003–0.050 J/item for baselines (190W). They are estimates, not energy measurements, and establish no simulation/hardware energy advantage. SO5 meters the bounded digital FPGA comparison; native physical oscillators remain a longer-term route. Beyond-encoding transfer and dynamics value remain open.

Public reproduction sources (checked 2 October 2026): v2 result, v2.1 ablations and stronger controls, versioned source and reproduction instructions, v3 frozen protocol, v3 result and v3 pinned source.

1.3.3 Methodology per objective and how it reaches the objectives in 36 months

SO1 / WP2 (M1–M18). Scale integrated/analytic phase representations to the SO1 target, then test relation algebra, NoRA and a ConceptARC subset re-encoded as relational scenes. Re-encoding is not native visual ARC evaluation; document modality/encoding losses. D2.1 fixes the reusable unit (relation coupling/motif and read-out), source training groups, validation selection and target composition groups. Freeze learned relation increments before held-out composition testing; no target labels or target-specific coupling fitting. Report frozen reuse separately from few-shot adaptation, whose examples, updated parameters and optimisation budget are shared with conventional controls. Relation renaming and topology/size shifts expose dependence on the encoding. Include the analytic phase map, a structured program control and adequately trained MLP/transformer/GNN and Scallop/Dolphin-class baselines [6,7], with equal source information and search budgets. Select the strongest conventional learned comparator on validation only. D2.1 freezes separate dynamics/noise endpoints and grouped analysis; the unchanged SO1 criterion governs MS1 at M12. D2.3 publishes the M18 result and adaptation decision.

SO2 / WP3 (M7–M30). Define the predicate language and read-out operators; build the constraint compiler; implement certificate generation (active-constraint sets, residuals, replayable integration seeds) and an independent checker that re-integrates and re-evaluates without trusting the producer. Replay validates a numerical witness; separate constraint/proof checks establish only the encoded logical properties. Every accepted result identifies its assumptions and check scope. Evaluate on causal (CLadder), constraint (ZebraLogic-class) and first-order (FOLIO) suites, reporting certificate rate, accuracy and J/reference-correct completed task.

SO3 / WP4 (M4–M34). Write the property catalogue (read-out correctness under bounded noise, monotone refusal, no acceptance without certificate, trace integrity) as a formal specification; verify the finite-precision read-out/acceptance/refusal logic by model checking and SMT; hash-chain retained reasoning traces against a trusted root. This detects covered trace modification, not factual error or every omitted event; the root and collection boundary are explicit. An independent team (subcontract) runs AgentDyn/AgentDojo-class attacks and targeted read-out attacks; a second subcontract audits the formal artefacts. D4.1 freezes eligible attack families, adaptive attack budget and utility sets before assessment. Attack-success rates include confidence intervals and family-level limits; the 2% target is not certification or security against untested attackers.

SO4 / WP5 (M13–M33). Formulate planning as control of slow variables with a differentiable model-predictive controller, reusing the contract discipline of the applicant's Petri-net controller compiler (scpn-control) for plan pre- and post-conditions; hierarchical levels correspond to time-scale separation; contingency basins and re-planning under perturbation. Plans are emitted in PDDL and validated; benchmarks: PlanBench Blocksworld and Mystery Blocksworld (600 instances each), including the ≥ 20-step subset on which the strongest reported reasoning model reaches 23.6 % 2.

SO5 / WP6 (M19–M36). Integrate on one metered node. Pilot A — structure discovery (Scientific Discovery): building on the applicant's phase-analysis toolkit (scpn-phase-orchestrator), whose grid-damping estimates are already validated against small-signal eigenvalues: given trajectories from an unseen coupled dynamical system (simulated oscillator networks, public power-grid frequency recordings from the Open Access Power-Grid Frequency Database (candidate OSF-linked recordings; eligibility and exact per-file terms established in the DMP, D1.1)), propose ranked coupling hypotheses and a next informative perturbation. Certificates check recorded numerical/constraint claims; causal recovery is scored only where a known generator or eligible intervention supplies ground truth. Real observational recordings test descriptive fit, uncertainty and abstention, without claiming identified causality. Pilot B — grounded decision support (DSS): answer bounded document-grounded questions over dated EU AI Act texts from EUR-Lex, subject to per-document reuse terms, with checked encoded reasoning steps and refusal when support is missing — a use case that other DeepRAP projects and TEFs can re-run on their own systems; built on the applicant's existing grounding and injection-detection components. Energy protocol: gross wall-plug joules for the complete local workflow divided by reference-correct, checker-admitted completed tasks, including feature extraction, language/phase computation, tools, checking, coordination and failed/retried attempts. Also report idle-adjusted energy and absolute batch energy; zero correct completions gives no finite efficiency claim. The ≤1.5kW inference envelope is commissioned at the wall, not inferred from PSU rating. Hosted frontier-service and off-site training energy are outside that local measurement; provider metering is separately identified, otherwise unknown. Cost/token/latency comparisons remain available without equating cloud energy to local energy.

Time plan feasibility. Existing integrator, probe, verification, grounding and coordination implementations reduce initial development work. Their compatibility, the typed compiler and releasable training pipeline still require integration effort. PI/SR1 own early software work within staffed allocations; ML/RSE start M7. M9 profiling precedes scaled training; the M12 gate leaves 24 months for the defined fallback. Compute commissioning must support these dates in the confirmed resource plan; existing prototypes do not prove deployment readiness.

1.3.4 High-risk issues and how they are addressed

Transfer to non-oscillator tasks remains uncertain and is gated at M12; certificate cost/yield and planning scalability require matched evaluations with independent checkers. The simulated substrate is costly and has no measured energy advantage. SO5 adds a bounded digital FPGA phase-kernel comparison using SC-NeuroCore's fixed-point/RTL/co-simulation toolchain. Existing pairwise Euler kernels do not implement the complete gated, staged or triadic KYMA architecture; that bridge is project work. Start with16/32 phase units, extending to64 only if fit/timing permits; SO1's128-unit software target remains separate. Freeze arithmetic, parameter import and read-out tolerances in WP2; prove declared overflow/wrap/admission properties in WP4; synthesize, validate and meter the board/host in WP6. Compare analytic phasor, dynamic and conventional structured kernels on matched tasks, including transfer overhead, failures and reference-correct coverage. Report bitstream/tool/part provenance, routed timing, latency and J/reference-correct task at kernel and host+board boundaries. No energy benefit is assumed. LLM training/inference remains on CPU/GPU. Native physical oscillators [14,15,19] remain longer-term context, not an unsecured fabrication dependency. An inadmissible hardware configuration stays disabled with a reproducible failure report; the digital pilots retain their CPU/GPU path.

1.3.5 AI robustness, interdisciplinarity, gender dimension, DNSH

Robustness. Technical robustness is an objective (SO3), not an afterthought: failure modes are reported, refusals are part of the output contract, and results are reproducible from published seeds and artefacts. Interdisciplinarity. The project combines nonlinear dynamics and synchronisation physics (substrate), computational neuroscience (oscillatory binding as the biological motivation), knowledge representation and neuro-symbolic AI (read-out and reasoning), formal methods (trust spine), and control theory (planning); each is represented in the team plan (Section 3.2) and in the independent evaluation. Gender dimension. The planned scientific datasets exclude identifiable human-subject data; any expert recruitment/feedback records require their own data-management review; sex/gender analysis of the research content is not applicable to the substrate research. It becomes relevant in Pilot B (decision support): the evaluation includes checks that refusal and answer quality do not differ systematically across gendered formulations of the same question, and datasets are screened for gendered bias, following the methods of the Commission's "Gendered Innovations 2" policy review (2020), referenced by the application template. DNSH. The main environmental factor is electricity for computation; it is measured and published per result, the metered deployment envelope is one node; separately budgeted training may run off-site and its resources are reported, and no hardware is procured beyond what the work requires.

1.3.6 Open science and research data management

Open science combines frozen protocols, immediate repository open access to publications, validation artefacts and rights-compatible replication packages. Publication fees are budgeted only for fully open-access venues; hybrid and printing fees are excluded. No-fee routes, including Open Research Europe where eligible, remain available. Software, model and data rights are assessed separately. Types and size: simulated trajectories and training logs (≈ 1–5 TB, compressed), benchmark task files (< 10 GB), model/adapter artefacts and retained run checkpoints (volume established by the profiled run/retention plan), energy logs. Findability: Zenodo DOIs per release, versioned citation files. Accessibility: open where rights permit, with justified privacy/security or third-party restrictions recorded in the DMP; no promise to redistribute restricted base weights. Interoperability: JSON/Parquet/HDF5 with documented schemas; PDDL for plans; ONNX where applicable. Reusability: explicit per-asset licenses, immutable revisions, seeds and environment locks. Curation: the applicant is responsible for data management; project experiment storage plus manifests/metadata; large checkpoints and datasets use owner-controlled storage with hashes and stable download links once host/retention/backup are resolved; repository and archive metadata remain appropriate to their licenses. A data management plan is delivered at M6 (D1.1) and maintained through phase-2 governance (D8.1). The rights matrix separates authored background, new contributions and third-party assets. Frontier-provider outputs are comparison evidence, not a presumed training/distillation corpus: any such reuse needs separately established rights. Reading access does not confer text/data-mining or training permission. Publicly accessible sources are screened for personal data; pretrained-base training provenance remains a disclosed limitation. Releases may provide adapters/manifests and lawful acquisition instructions where base weights or data cannot be redistributed.

1.3.7 Specialised-model learning and controlled integration

Common model and tool architecture

Inputs comprise language, document spans and numerical observations encoded by a documented feature pipeline. Pilot A raw signals are processed by domain tools into features, units, uncertainties and relations; text tokenisation is not presumed to understand raw waveforms. The language model proposes a typed task graph. The learned oscillator module forms or composes relation motifs; a symbolic compiler produces explicit predicates, constraints or plans. Independent numerical/constraint checks validate only their declared properties. The response renderer cites checked outputs and records refusal when admission conditions fail.

Typed envelope fields: task/source/revision IDs; observation feature schema and units; model/tokenizer/adapter versions; input predicates and uncertainty; requested tool and bounded parameters; relation/motif references; proposed plan; checked properties and verifier version; result status; evidence hashes; resource-use record. Initial transport may use existing Synapse interfaces; interface compatibility is part of the proposed integration work.

LLM training starts with supervised examples of grounded typed outputs and tool selection. A candidate objective combines token cross-entropy, supervised predicate/read-out consistency and coupling regularisation. The symbolic checker is outside the gradient path. Staged training of the language interface and oscillator couplings is the default; a jointly differentiable coupling/read-out branch requires actual interface feasibility and is an ablation. Loss coefficients, representation dimensions and optimisation/search budgets are frozen on validation data in D2.4/D3.1 before confirmatory training. LoRA/QLoRA supplies an established efficiency technique. KYMA tests the learned phase interface; its formal checker remains non-differentiable.

Apertus is a current Swiss candidate, not a fixed selection. Candidate selection requires exact base/tokenizer revisions, training-stack feasibility, license/access/AUP review and redistribution terms. Selection prioritizes one technically suitable European-region base and an independent comparator; no provider, institution or HPC allocation is presumed committed. No scratch foundation-model pretraining or initial 70B training is budgeted. Continued domain pretraining or reward-based training requires a separate M12 justified branch and a bounded reallocation, not an automatic new commitment.

Corpus and ground truth

Planning target per pilot:10,000 training cases,2,000 validation cases and2,000 held-out test cases. M6 establishes a source/task inventory and small rights-cleared samples, not a claim that the full corpora already exist. M9 assesses eligible group counts, label effort and model feasibility before scaled collection/training. Cases may include multiple eligible observations/documents, solver results, plans and failure examples. Split source/generator/topology/composition groups before creating prompt variants. Each task-source/generator/composition group stays in one split. For Pilot B, immutable public reference texts may be a common retrieval context in all splits; task profiles, answer labels and rule-composition templates are separated. The primary Pilot B claim is conditional on that fixed corpus and concerns unseen profile/composition groups, not transfer to independent legal corpora. Keep sealed evaluation questions, gold answers and labels out of training and retrieval indexes. Evaluation may retrieve permitted held-out source documents as task context; document availability is identical for compared systems, with no training on their held-out source groups. Upstream model contamination cannot be ruled out solely by our split; add newly generated hidden cases and disclose the limitation.

Pilot A supervision combines project-controlled simulations and eligible open recordings. Ground truth distinguishes simulator-known causal structure from hypotheses inferred from observational recordings. Interventions or known generators supply causal checks where applicable; descriptive synchronization alone is not causal proof. Pilot B supervision combines eligible EU documents, expert-reviewed constraint encodings, answer spans and independently validated plans. Document interpretation and factual support remain separately assessed from formal constraint consistency. Pilot B targets explicit document spans, definitions and bounded synthetic profiles. Its reasoning strata include multi-step joins of explicit conditions, version/date distinctions and insufficient/conflicting-support cases; source spans and intermediate predicates support independent checking. The protocol distinguishes retrieval-only items from composition tasks and includes a conventional rule/tool baseline. If that baseline reaches ceiling, this does not establish a phase-representation quality advantage. Unresolved jurisdictional or interpretive questions are excluded from the confirmatory endpoint or refused. It is not legal advice or a compliance determination.

Each pilot's planned 2,000 test cases form four disjoint 500-case cohorts: WP3 M24 diagnostics/M30 comparison and WP6 M30 diagnostics/M34 confirmation. Group separation applies across cohorts. Provider campaigns name the cohort manifest and share cases for paired comparison; five frontier sampling repeats are distinct from five training seeds and do not add independent cases. Before M34, freeze model/adapters, provider revisions, decoding, tools, checker and criticism policy; lock primary predictions before criticism. Gold answers/diagnostic feedback never tune the primary model or guide its search. Cloud transmission requires established access/retention terms. Profiling uses non-test cases. Opening a cohort retires it from future confirmation; the first three remain exploratory unless separately validated under their own frozen endpoints and cannot be pooled into the final claim.

M6 inventory measures source/generator/topology diversity and annotation time; M9 checks feasibility and grouped precision. Five hundred cases need not be 500 independent groups: template variants do not add independence. If capacity is insufficient, reduce auxiliary training volume, revise the design before freeze, or report an inconclusive endpoint; preserve a sufficiently powered primary test rather than dilute it.

Training examples may use controlled generators and checked encodings, with expert source/rule review and sampled quality checks; they are not all individually expert-annotated. Annotation effort and reference quality are checked at M6/M9. Reference definitions and independent system assessment remain separate. Cases retain source/licence, tool revision, units, seeds, split and checker records, including contradictions, refusals and failures.

Controlled comparisons and run budget

Compare frozen prompted LLM; the same model with retrieval/conventional tools; plain specialised LLM; conventional structured specialisation with the same checker; and the KYMA hybrid with the same checker. Report a classical true-domain planner as an explicitly privileged upper bound where appropriate. The plain specialised-model comparison may reuse conventional specialisation weights with the checker removed. Hyperparameter searches, context/response limits, source access and allowed tools are matched or their differences documented.

Primary training inventory: two pilots × two trainable conditions × five seeds = 20 runs on the first base. A second base adds 20, subject to feasibility. Up to 10 targeted training ablations produce a 50-run ceiling. Frozen-model, checker removal, retrieval and memory comparisons need additional inference evaluation but do not all require fresh training. Unique run IDs distinguish training, feasibility, transfer and evaluation workloads.

The 50-run ceiling has an explicit planning allocation: first-base 20 runs at up to 400 accelerator-hours each and conditional second-base 20 at up to 250 each fit the 13,000-hour WP3 envelope; up to 10 transfer/ablation runs at 500 each fit WP5's 5,000 hours. The WP2 3,000-hour feasibility and WP6 4,000-hour evaluation envelopes complete the 25,000-hour total. These are proposed runtime ceilings, not demonstrated throughput. Within each base, hybrid and conventional conditions receive the same profiled budgets; the second base is secondary and proceeds only if feasible without compromising the first-base primary evaluation.

Profile separate rights-cleared non-test samples at bounded contexts; request IDs never expose the four pilot cohorts or sealed M34 labels. The provisional cap is 100M processed tokens/run including repetitions; freeze sequence lengths, training tokens and accelerator type/count after profiling. The envelopes above are estimates. Record node-hours separately from accelerator-hours; small GPUs do not automatically pool memory.

Primary endpoint and inference. Valid-task coverage is reference-correct, checker-admitted completion divided by all eligible test tasks; refusals, timeouts and failed calls remain in the denominator. Each final 500-case pilot cohort has independently reviewed references and frozen tolerances before predictions are opened.

Pilot Required output and reference What counts as a primary completion
A: structure discovery A submitted coupling hypothesis scored against a held-out known generator; prescribed numerical/constraint checks Structural correctness within frozen edge/parameter tolerances and a valid scoped witness. Ranked alternatives cannot hide a wrong committed answer. Observational fit and proposed next-perturbation utility are separate diagnostics
B: document support Answer, dated source spans and encoded derivation, scored against reviewed answer/support labels Correct supported answer and valid derivation. A consistent derivation of an unsupported statement fails; refusal of a solvable item reduces coverage

Select the strongest conventional comparator on validation only. Use paired source/topology/composition groups, seed-level effects and the frozen hierarchical analysis; five seeds are not independent datasets. SO6's joint claim requires both pilots to meet their unchanged 10 pp point target and adjusted interval criterion. Success in one pilot supports only that domain; a pooled mean cannot establish joint transfer. D3.1 fixes adjudication, abstention and missing-reference rules. Reference defects are reviewed blind to system identity with an auditable common-case disposition, never removed selectively after outcomes. The external system assessor replays scoring and coverage from frozen predictions; reference authorship and independent assessment are disclosed separately. Validity, semantic correctness and risk–coverage curves accompany the primary endpoint.

Precision and target attainment. D3.1 freezes the grouped paired method, the independent-group requirement and both-pilot success rule before confirmatory training. Illustratively, for independent paired binary outcomes and a true 10 pp difference, a two-sided Bonferroni planning bound for two endpoints gives roughly 181, 371 or 561 independent observations for 80% power when discordance is 0.2, 0.4 or 0.6. These analytic sensitivity values do not establish power for clustered data or for the joint claim. At the true 10 pp effect, the separate point-estimate ≥ 10 pp rule passes only about half the time; a confidence interval excluding zero is not a guarantee of target attainment. For illustration, a true 15 pp effect with discordance 0.4 requires approximately 248 independent observations for 90% probability per pilot of also exceeding the 10 pp point target under a normal approximation. Two such per-pilot probabilities would bound joint attainment at about 80% without assuming independence. With 500 cases grouped as 50 groups of 10 and within-group correlation 0.05, the illustrative effective size is only about 345. M9 therefore checks actual source/generator diversity and plausible paired effects/correlation, using simulation of the frozen grouped analysis; no fixed group count is asserted as already available. If the unchanged 500-case final allocation cannot support the agreed precision, revise the design and workload openly before freezing or report the endpoint as inconclusive. Five training seeds do not multiply independent evidence.

Metrics include held-out-composition correctness and valid-task coverage; unsupported acceptance; refusal and utility; independently valid certificate/plan rates; few-shot/sample-efficiency curves; latency/tokens/tool calls; training accelerator-hours; wall-plug inference joules per correctly completed task; and actual monetary cost. Off-site energy uses provider metering where available, otherwise a clearly labelled estimate/unknown. Publish training and inference cost separately and amortize training only with an explicit query-volume assumption.


Intended impact, open research and collaboration

2.1 Potential impact

(a) Unique contribution.

To the outcomes of the Challenge. KYMA proposes a contribution to a future DeepRAP portfolio an experimentally distinguishable contribution: a physics-inspired substrate (NIF) with formal guarantees on its trust-critical parts (FMI/TM), with the precise proof assumptions and empirical limits disclosed. Concretely, the project proposes five outputs that map one-to-one onto the Challenge's expected outcomes:

Expected outcome (WP 2026) KYMA output that serves it Usable by others without KYMA's substrate?
Architectures trainable and deployable with constrained resources Integrated system on one metered node; energy and cost protocol (D6.2) Protocol: yes
Provable trustworthiness mechanisms, AI Act alignment Verified read-out/acceptance/refusal components; certificate format and independent checker (D3.2, D4.2); AI Act assessment (D1.3) Checker and certificate format: yes
Cognitive AI system at TRL 4 on real tasks Pilots A (structure discovery) and B (bounded document-grounded decision support) (D6.3, D6.4) Pilot B task set: yes
New methods and metrics for reasoning, trustworthiness and resource use Composition gap, certificate rate, reference-correct valid plans per joule, energy per reference-correct completion Yes
FAIR; synergies with TEFs, EBRAINS, RAISE, AIoD, Quantum Flagship Open releases; offers listed in 2.3 Yes

To the wider impacts. If the substrate hypothesis holds, Europe gains a line of reasoning systems whose conclusions have an inspectable phase state and a scoped verifiable read-out — providing testable evidence relevant to transparency, robustness and human oversight, without claiming legal compliance by architecture alone. Compatibility with analogue/neuromorphic implementations is tested rather than presumed. The bounded model programme contributes rights-compatible locally deployable specialisations, reproducible training recipes and checked-task benchmarks, prioritizing an eligible European-region base. Frontier services are comparison instruments; buying access is not itself European model development or compute sovereignty. If transfer fails, documented negative results and scoped checker/resource artefacts remain useful, without satisfying the central breakthrough claim.

Target groups, specified. (1) Research groups in neuro-symbolic AI, dynamical-systems machine learning and neuromorphic engineering — they receive toolkits, protocols committed before the run and benchmark families; (2) developers of AI for regulated settings — SMEs and public bodies that must document how a system reached a conclusion; they receive the certificate checker and the bounded document-support pilot; (3) neuromorphic and analogue hardware developers — they receive cognitive workloads and a motif-to-hardware mapping beyond classification benchmarks; (4) testing, certification and standardisation actors — the sectoral AI Testing and Experimentation Facilities (TEF-Health, AI-MATTERS, CitCom.ai, agrifoodTEF) and the CEN-CENELEC JTC 21 work on AI Act harmonised standards (including logging and risk-management work, with applicable work items rechecked before contribution; Swiss participation through the SNV mirror committee); they receive measurable reasoning-trust and resource-use metrics.

Potential negative effects. Compute energy (measured and capped; Section 1.3.5); misuse of red-team findings (published as findings and defences, not as attack kits); over-trust in certified answers (certificates state what was checked and what was not).

(b) Scale and significance (assumptions stated; one method for all estimates).

Effect Baseline KYMA estimate at M36 Assumption / how measured
Uptake of KYMA artefacts by other groups 0 ≥ 3 external groups, of which ≥ 1 DeepRAP portfolio project, run a KYMA benchmark family, the certificate checker or the energy protocol Verified external execution with versioned outputs; forks and citations are reach indicators, not evidence of use
Share of reasoning results reported with energy per solved task in the DeepRAP benchmark goal-level metric already proposed [24]; uptake of this common-node protocol not established propose energy per solved task as an optional benchmark axis; adoption is a portfolio decision Decision of WG1; KYMA provides scripts and reference measurements
Cost of checked completion in Pilot B Conventional checked specialisation on the same frozen Pilot B cases, sources, tools and node; no baseline value yet measured Report J and EUR per reference-correct completion with coverage and certificate rate D6.2/D6.3 matched task-level measurements; ARC and AI Energy Score remain methodological context
Hardware relevance Oscillator hardware trained on MNIST-class tasks and Max-SAT [14,15] ≥ 1 trained KYMA motif set mapped to an oscillator-hardware model with the accuracy loss reported WP6, using published device models; no own hardware claims

Route to use. By M18, WP7 provides a replayable typed task, checker and rights manifest to prospective dynamical-ML and neuro-symbolic users. By M30, the integration candidate adds failure/refusal and resource records; M36 releases include a clean-environment reproduction command and the independent assessment limits. PI/AREP own scientific handover, RSE the runnable package and LIAISON access follow-through. Log invitations, execution support and versioned user outputs; the three-group uptake target is met only by external replay, not a meeting or repository fork. Non-response does not prevent the applicant-controlled releases. Certification-support and hardware licensing remain conditional on demonstrated properties, third-party rights and user demand; no sales or market-size forecast is claimed.

(c) Requirements and barriers beyond the project. (i) Hardware: native oscillator hardware is at research stage; KYMA's trust spine, metrics and pilots are useful in digital form, and the hardware mapping is designed to be picked up by hardware projects (e.g. through the EBRAINS neuromorphic platforms SpiNNaker and BrainScaleS, which advertise remote access for evaluation/basic research; task suitability, allocation and any phase-to-spiking mapping require separate validation and are not committed KYMA resources). (ii) Regulation: harmonised standards under the AI Act are still being developed; KYMA offers metrics to that process rather than assuming its outcome. (iii) Adoption: certified reasoning means users accept refusals; Pilot B measures refusal rate and utility together. (iv) Talent: the profiles are scarce; see 3.1. Project implementation remains conditional on funding, recruitment, eligible resources and independently validated capabilities. No external facility allocation or partner commitment is assumed.

2.2 Innovation potential

Proof of principle. The TRL-4 demonstration (MS5) is the proof of principle: two pilots combining eligible real observations/documents with reference-labelled simulations and bounded document tasks, one metered node, two independent evaluation reports. The MS1 gate and the pre-committed fallback make the proof of principle realistic even if the substrate scales less than hoped: the checked conventional fallback is evaluated against the same integration criteria; TRL 4 is reported only if the demonstrated laboratory system meets them.

Background software, model weights and datasets have separate licences. Proposed exploitation routes include commercial software licensing, certification-support services and bounded hardware mappings, subject to demonstrated properties and third-party rights. The programme makes no claim of legal certification, future grant eligibility or secured customers.

Intellectual property. Background: the applicant’s authored contributions to six systems (Section 3.2), with third-party dependencies, judges, pretrained weights and corpora recorded separately. Background ownership is not a claim to all incorporated assets. Foreground candidates, subject to contribution and third-party rights: trained coupling topologies, the constraint compiler, certificate formats, the energy protocol, specialised model adapters, typed language interfaces and rights-cleared task corpora. Measures: patentability screening at MS1 and MS3 before publication (priority filing only if a claim survives prior-art search — AKOrN, KoPE and related work are prior art for generic oscillator reasoning); release the training recipes needed to reproduce scientific claims; retain only justified non-essential confidential material under the DMP/IP plan; copyright and dual licensing for software; trademark check for "KYMA". Regulation, certification and standardisation are assessed in D1.3: mapping of KYMA outputs to AI Act requirements (transparency, accuracy/robustness, human oversight) and to the relevant harmonised-standard work items (logging and risk-management work; quality-management standards as context). Adoption of a standard is distinct from completed harmonisation/OJ citation and does not itself establish a presumption of conformity.

Proposed researchers would lead scientific work according to their expertise and receive authorship credit for actual contributions. Recruitment and external collaboration remain conditional on funding and separate agreements. No institution, provider, position or portfolio appointment is confirmed.

2.3 Communication and Dissemination

Target group Main message Tools and channels Measure of success (by M36)
Research community Whether an oscillator substrate can carry compositional reasoning — with the evidence, positive or negative Open-access papers and preprints (targets: the NeSy conference, NeurIPS and ICLR workshops, and journals such as Neuromorphic Computing and Engineering, Neural Computation and Chaos; open-access fees are budgeted in WP3 and WP4); protocols committed before the run; replication packages ≥ 4 open-access papers or preprints; every result with a public replication package
DeepRAP portfolio Shared benchmark families, certificate checker, energy protocol WG1/WG3 contributions, annual portfolio meeting, joint-pilot offer Deliver ≥ 2 complete portfolio integration/benchmark contribution packets and 1 joint-pilot-ready kit; seek co-leadership and external pilot uptake
Regulated-AI developers, SMEs An encoded reasoning step can be checked and costed Technical notes, two open webinars, Enterprise Europe Network multipliers ≥ 2 webinars; ≥ 20 organisations reached
Hardware developers Cognitive workloads for oscillator/neuromorphic hardware Targeted workshop with oscillator-hardware groups, EBRAINS community channels 1 open technical workshop and 1 reproducible motif-testing package; seek testing by an external hardware group
Policy and standardisation Measurable reasoning-trust and resource-use metrics Input to AI Act implementation consultations and to relevant logging-standard enquiries via the SNV mirror committee; contacts with TEF-Health and AI-MATTERS ≥ 1 formal contribution to a consultation or standardisation work item
General public Why it matters how an AI reaches a conclusion — explained with oscillators Plain-language project page, one public talk per year, short videos of the pilots 3 public talks; project page kept current

Output accountability. Applicant-controlled packages, releases and events are deliverables. External participation, co-leadership, adoption and peer-review acceptance are uptake targets; refusals or non-response are recorded and do not prevent delivery of the standalone digital research programme.

Timing. Communication starts at M1 (project page, first protocol commitment before the run), dissemination of results from M12 (MS1 report, whatever its outcome), exploitation activities from M18. Plan. D1.2 is delivered at M6 and updated through M18; phase-2 implementation and the M36 update are included in D8.2. Policy feedback. D1.3 in WP8 (M20, updated M34) summarises what KYMA's measurements imply for assessing reasoning trustworthiness and resource use, for use in AI Act implementation discussions. All claims follow the project's evidence rule: no result is communicated without its protocol, its baseline and its cost.


Proposed work packages, timeline, outputs and risks

Proposed roles: PI = principal investigator; SR1/SR2 = senior researchers; ML = machine-learning specialist; RSE = research software engineer; LIAISON = technical liaison; AREP = academic and reproduction support; ADMIN = administration. Effort is proposed, with no funded appointments implied.

Figure 2 — revised work plan and decision gates.
Figure 2 — revised work plan and decision gates. All months are relative to a future funded start.
Open the full-size figure

Dependencies: M6profile → M9training feasibility → M12substrate go/adapt → M24specialisation/checker/custody → M30comparison → M33transfer/planning → M34pilot/independent reports → M36TRL4/release. Eight WP windows are specified in Table 3.1a.

WP M1–6 M7–12 M13–18 M19–24 M25–30 M31–36
WP1 Management phase1 planned planned planned
WP8 Management phase2 planned planned planned
WP2 Substrate & abstraction (SO1) planned planned MS1 planned
WP3 Motif–symbol & certificates (SO2) planned planned planned MS2 planned
WP4 Verified trust spine (SO3) planned planned planned planned MS3 planned planned
WP5 Planning (SO4) planned planned planned planned MS4
WP6 Integration & TRL-4 (SO5) planned planned planned MS5
WP7 Portfolio activities planned planned planned planned planned planned

Inter-relations (arrows in Figure 2). WP2 → WP3 (motifs to read out), WP2 → WP5 (substrate to control), WP3 → WP4 (read-out to verify), WP2 → WP3 (specialised training), WP3 → WP5 (typed plans/transfer), WP3/WP4/WP5 → WP6 (integration), WP6 → WP7 (benchmark and pilots to the portfolio), WP1/WP8 provide continuous management across the two phases.

Table 3.1a — List of work packages

WP Title Lead Start End PM
WP1 Management, provenance and recruitment — phase1 Participant 1, PI 1 18 13
WP2 Trainable oscillator abstraction and training feasibility Participant 1, SR1 1 18 39
WP3 Language–motif–symbol interface and specialised model training Participant 1, SR2 7 30 30
WP4 Verified acceptance and independent system evaluation Participant 1, SR2 4 34 21
WP5 Checked planning and cross-task transfer Participant 1, SR1 13 33 29
WP6 Integrated pilots, model release and resource measurements Participant 1, RSE 19 36 34.25
WP7 DeepRAP interoperability and portfolio activities Participant 1, PI 1 36 17
WP8 Management, provenance and dissemination — phase2 Participant 1, PI 19 36 14

Table 3.1b — Work package descriptions

WP1 / WP8 — Management, provenance and dissemination in two phases

Both phases cover project management/reporting, separated internal/external administration, data management, dissemination/exploitation, regulatory/IP review, model/data rights, procurement and contribution provenance, staffing and project-specific access coordination. PI retains decisions, ADMIN prepares records, the fiduciary performs bookkeeping/payroll and LIAISON coordinates technical access/logistics. General ecosystem sales and future-institute fundraising are outside the action.

WP1,M1–M18: PI2,ADMIN5,LIAISON6=13 PM. Complete initial employer/procurement/rights arrangements, recruitment records and first-period implementation controls. Outputs: D1.1/D1.2/D1.4 at M6, updates through M18 and D1.5 first-period management report at M18. Grant-contingent selection preparation supports the M1 liaison start; the six-week shortlist checkpoint must precede the relevant start, rather than begin only at M1. Later checkpoints support M4/M5/M7 research appointments. Agreed start dates, funding availability, employer arrangements, work authorization and fiduciary scope precede paid employment. Completed records and reports define phase-1 completion.

WP8,M19–M36: PI3,ADMIN5,LIAISON6=14 PM. Continue distinct second-period controls, provenance/rights maintenance, staff and procurement monitoring, policy/IP assessment and final dissemination/reporting. Outputs: D1.3 at M20 (updated M34), D8.1 updated governance/DMP/provenance packet at M30 and D8.2 final implementation/D&E/management report at M36. Subsequent updates do not retrospectively reopen WP1. Neither phase depends on an institute or grant transfer.

WP2 — Trainable oscillator abstraction and training feasibility (M1–M18)

Tasks: T2.1 gated integrator; T2.2 three non-oscillator task families; T2.3 matched and well-trained baselines, including phase-arithmetic controls; T2.4 five-seed MS1 evaluation; T2.5 replication; T2.6 pretrained-base selection and typed numerical/relation schema; T2.7 corpus sampling, grouped splits and loss/training feasibility; T2.8 selected phase-kernel quantization and software/RTL interface contract. AREP supports scientific requirements/source/task review and reproduction preparation within WP2; SR1 supplies 15 PM in M4–M18; PI/ML/RSE cover early interface and planning prerequisites. Effort: PI 6, SR1 15, RSE 4, LIAISON 8, ML 4, AREP 2 = 39 PM.

Outputs: D2.1 protocol at M4, D2.2 toolkit at M9, D2.3 results/replication at M18, D2.4 training-feasibility packet at M9. PI/SR1 own M6 sample/specification readiness; ML profiles after its M7 start. The M9 gate rejects infeasible rights/stack/resource assumptions before scaled procurement; MS1 at M12 decides the substrate branch.

WP3 — Language–motif–symbol interface and specialised model training (M7–M30)

Tasks: T3.1 predicate/read-out design; T3.2 constraint compiler; T3.3 independent certificates; T3.4 causal/constraint/first-order evaluation; T3.5 two supervised language/typed-relation specialisations; T3.6 staged coupling training and conventional-representation comparator; T3.7 versioned model registry and sealed held-out comparison. ML implements data/training pipelines under SR2's specification, SR1's abstraction input and PI oversight. Effort: PI 4, SR1 6, SR2 14, ML 6 = 30 PM.

Outputs: D3.1 specification/protocol at M12, D3.2 checker at M22, D3.3 reasoning report at M30, D3.4 specialisation/training manifests at M24 and D3.5 matched comparison at M30. Training failures produce a documented conventional-interface fallback and a valid failure report. RSE receives maintenance/reproduction custody by M24 before the ML term ends.

WP4 — Verified acceptance and independent system evaluation (M4–M34)

Tasks: T4.1 formal property catalogue; T4.2 acceptance/refusal verification; T4.3 independent robustness/system evaluation, including primary-scoring/reference/coverage replay; T4.4 separate formal-artefact audit; T4.5 admitted FPGA arithmetic/overflow/wrap and interface properties. The catalogue covers model/tool envelopes, schema rejection, source binding, numerical assumptions and checked-plan admission. Statistical model confidence is not a proof. Independent robustness and formal providers would be selected for suitability, independence and best value; no provider is currently committed. Effort: PI 3, SR2 13, RSE 5 = 21 PM.

Outputs: D4.1 catalogue/plan at M10, D4.2 formal artefacts and verification decision at M24, D4.3 independent integrated-evaluation report at M34. Each assessed model/checker revision is explicit; deployment-critical counterexamples trigger redesign or disablement. Findings include reproducible scope and unresolved limits.

WP5 — Checked planning and cross-task transfer (M13–M33)

Tasks: T5.1 hierarchical slow-variable control; T5.2 contingency/replanning; T5.3 learned transition models and PDDL admission checks; T5.4 matched classical/frontier evaluation; T5.5 specialised typed tool/plan proposals; T5.6 held-out/few-shot transfer; T5.7 oscillator/conventional representation and optional memory ablations. All compared pipelines use the same source/tool/validator budget. Effort: PI 3, SR1 8, RSE 6, LIAISON 6, ML 4, AREP 2 = 29 PM. AREP coordinates scientific reproduction and transfer requirements within WP5. PI/ML/RSE own M13–M18 prerequisites while SR1 is fully assigned to WP2; SR1's intensive planning work starts after M18.

Outputs: D5.1 protocol at M18, D5.2 planning results and D5.3 transfer/sample-efficiency report at M33. D5.1 names source/target tasks, reusable parameters and the permitted adaptation budget; report reused, reinitialised and newly trained parts separately. Two pilot successes alone do not establish cross-domain reuse of the same learned motif. Visits gather relevant requirements and test evidence. A fusion/neuromorphic case is optional within existing resources and rights, not a third pilot or real-device actuation.

WP6 — Integrated pilots, model release and resource measurements (M19–M36)

Tasks: T6.1 integrated metered system; T6.2 energy/cost measurement; T6.3 the two pilots; T6.4 independent TRL4 evaluation; T6.5 adapter/inference integration, custody, compatibility and releasable model/data packages. Provenance-bound tools/memory and separate training/inference accounting are included. AREP contributes technical pilot reviews/test-session preparation within WP6. T6.6 SC-NeuroCore FPGA phase-kernel implementation comparison: synthesis readiness M24, board bring-up M27, measurements M30–31 and integrated results M34. Within existing effort, reserve RSE0.5 PM in WP2 (M9–10),1.5 PM in WP4 (M12–14),6 PM in WP6 (0.75 each M24–31), and SR2 1 PM in WP4 (M21–22); monthly capacity is retained. Effort: PI 4, SR1 2, SR2 3, RSE 13, LIAISON 6, ML 4, AREP 2.25 = 34.25 PM.

Outputs: D6.1 platform at M27, D6.2 resource protocol at M24, D6.3 pilot reports at M34, D6.4 validation and D6.5 replication package at M36. One local node is the deployment target; permitted off-site training is separately budgeted. The digital pilots can be completed without unsecured external facilities. TRL4 validation requires a frozen integrated model/checker/tool revision, replayable end-to-end runs on the commissioned node, reviewed references and retained failure/refusal/resource records in both pilots, plus independent assessment of the implemented capabilities and their limits. A simulation score or a report of negative findings alone does not establish TRL4. Native oscillator hardware fabrication or a third pilot is not on the contractual critical path.

WP7 — DeepRAP interoperability and portfolio activities (M1–M36)

Tasks: portfolio governance and working groups, strategic-plan contributions, benchmark co-creation, licensed task/model artefacts, typed result interfaces and resource metrics. The proposed PI would represent KYMA in WG1 and WG3; LIAISON supports WG2 and WG4 and coordinates attendance under PI responsibility. Plan approximately quarterly online working-group meetings and three annual in-person portfolio meetings over the action, with exact dates fixed by the portfolio. M6 specifies reusable benchmark/certificate/resource envelopes; M18 supplies the first replayable example and licensing manifest, M30 a cross-project integration candidate, and M36 reference releases. Chair and joint-pilot uptake remain proposals, not appointed roles or secured participation. LIAISON supports access and follow-through; AREP contributes scientific benchmark/interface/meeting packets; PI owns scientific portfolio decisions. Effort: PI 5, SR1 2, SR2 2, RSE 2, LIAISON 4, AREP 2 = 17 PM. D7.1 reports at M18/M36 record the contributions. Joint work is offered; uptake and partner commitments are not assumed.

Table 3.1c — Deliverables

No Name WP Type Dissemination Month
D1.1 Data management plan 1 DMP PU 6
D1.2 Plan for dissemination and exploitation incl. communication 1 R SEN 6
D1.3 AI Act, IP and standardisation assessment 8 R PU 20
D2.1 SO1 protocol committed before the run and task families 2 R + DATA PU 4
D2.2 Scalable gated-coupling toolkit release 2 OTHER (software) PU 9
D2.3 SO1 results, MS1 decision and replication package 2 R + DATA PU 18
D3.1 Predicate language, read-out specification, SO2 protocol 3 R PU 12
D3.2 Certificate format and independent checker 3 OTHER (software) PU 22
D3.3 Certified-reasoning results on causal, constraint and first-order suites 3 R + DATA PU 30
D4.1 Property catalogue and verification plan 4 R PU 10
D4.2 Verified trust-spine components and proofs 4 OTHER PU 24
D4.3 Robustness results and independent red-team report 4 R PU 34
D5.1 Planning architecture and SO4 protocol 5 R PU 18
D5.2 Planning benchmark results with validated plans 5 R + DATA PU 33
D6.1 Integrated platform on the metered node 6 DEM PU 27
D6.2 Energy and cost reporting protocol 6 R + OTHER PU 24
D6.3 Pilot A/B and FPGA implementation-comparison reports 6 R PU 34
D6.4 TRL-4 validation report 6 R PU 36
D7.1 Report on portfolio activities (one per reporting period) 7 R SEN 18, 36
D1.4 Model/data rights and provenance register 1 R SEN/public summary 6
D2.4 Specialised-model training feasibility 2 R+DATA PU where permitted 9
D3.4 Two specialisations and training manifests 3 OTHER+DATA PU where licensed 24
D3.5 Matched specialisation/abstraction comparison 3 R+DATA PU 30
D5.3 Transfer and sample-efficiency report 5 R+DATA PU 33
D6.5 Reproducible model/data/deployment package 6 OTHER+DATA PU where licensed 36
D1.5 First-period management and implementation report 1 R SEN 18
D8.1 Updated governance, DMP and provenance packet 8 R SEN/public summary 30
D8.2 Final implementation, dissemination and management report 8 R SEN 36

Table 3.1d — Milestones

No Name WPs Month Means of verification
MS1 Substrate go/adapt gate 2 12 Gate report against the D2.1 criterion (≥ 10 pp in ≥ 2 of 3 families); decision recorded, fallback activated if not met
MS2 Certified read-out operational 3 24 All accepted outputs pass the specified checker; certificate-yield target ≥ 80 % of the frozen reference-solvable set; semantic correctness separately reported
MS3 Trust-spine verification decision 4 24 Property catalogue resolved; critical admission properties proven for any admitted configuration, otherwise redesign/disable gate; artefacts and limits public
MS4 Planning gate 5 33 D5.2 frozen comparison complete; all executed plans validated and SO4 coverage gain/refusals reported. Pass/adapt decision records whether the unchanged SO4 target was met; validation alone does not establish planning superiority
MS5 TRL-4 demonstration 6 36 Two pilots run end-to-end on the metered node; independent evaluation reports attached
MS6 Training feasibility decision 2 9 Rights-cleared corpus sample, typed interface, profiled runtime and bounded run plan
MS7 Specialisation/custody handover 3,6 24 Model packages or justified failure/fallback report, reproducible setup and retained run/checker manifests; RSE receives maintenance/reproduction custody using its WP6 effort. This is readiness, not the final SO6 endpoint

Table 3.1e — Critical risks for implementation

Risk Likelihood / severity WPs Mitigation
Substrate advantage does not transfer to non-oscillator tasks Medium–high / high 2, 3, 5 Tested first; MS1 at M12; pre-committed fallback (modules inside conventional pipelines); SO3 and energy protocol independent of the outcome
Benchmark and baseline drift (field moves fast) High / low–medium 2–5 Relative criteria against matched baselines; baselines re-verified at each protocol commitment before the run
Hardware timing or energy envelope fails Medium / medium–high 2,6 Commission against real workload and wall-plug limit; competent early setup, M9 interface/M12 phase-board gate; revise the resource plan. Preserve CPU/GPU evaluation and report actual energy/coverage; no automatic TRL4 claim
Independent evaluators unavailable Low / medium 4 Identify suitable independent providers early; competitive scope/selection by M12, staged verification at M24 and integrated evaluation by M34; no provider already committed
Model/data rights or training feasibility fail Medium/high 1–3 Alternate eligible base, rights-cleared corpus, M6 profile/M9 gate; publish limits and preserve checked conventional fallback
Training regresses or transfer does not hold Medium/high 3,5 Frozen/conventional matched baselines, sealed group splits and negative-result publication

References

[1] ARC Prize, ARC-AGI leaderboard and benchmark descriptions, https://arcprize.org/leaderboard (accessed 01.10.2026); freeze benchmark/harness revisions per evaluation. Primary source.

[2] Valmeekam et al., arXiv:2409.13373v1 (2024). Primary source.

[3] Xie et al., TravelPlanner, ICML 2024. Primary source.

[4] Hao et al., Large Language Models Can Solve Real-World Planning Rigorously with Formal Verification Tools, arXiv:2404.11891v3 (2025). Primary source.

[5] Lin et al., ZebraLogic, ICML 2025. Primary source.

[6] Li et al., Scallop, PLDI 2023. Primary source.

[7] Naik et al., Dolphin, ICML 2025; Biberstein et al., Lobster: A GPU-Accelerated Framework for Neurosymbolic Programming, ASPLOS 2026, DOI 10.1145/3760250.3762232. Primary source.

[8] Jiao et al., DeepProofLog: Efficient Proving in Deep Stochastic Logic Programs, AAAI 2026, DOI 10.1609/aaai.v40i27.39396. Primary source.

[9] Kamali et al., NeSyCoCo: A Neuro-Symbolic Concept Composer for Compositional Generalization, AAAI 2025, DOI 10.1609/aaai.v39i4.32439. Primary source.

[10] Lake & Baroni, Nature 623 (2023). Primary source.

[11] Das et al., When No Paths Lead to Rome: Benchmarking Systematic Neural Relational Reasoning (NoRA), NeurIPS 2025 D&B, arXiv:2510.23532v1. Primary source.

[12] Löwe et al., TMLR 2022; Löwe et al., NeurIPS 2023; Muzellec et al., arXiv:2502.21077. Primary source.

[13] Rusch & Mishra, arXiv:2010.00951; Rusch & Rus, ICLR 2025; Effenberger et al., PNAS 122 (2025). Primary source.

[14] Rageau & Grollier, Neuromorph. Comput. Eng. 5 (2025); Laydevant et al., Nat. Commun. 15 (2024). Rageau & Grollier preprint.

[15] Delacour et al., arXiv:2505.07179. Primary source.

[16] Miyato et al., AKOrN, ICLR 2025. Primary source.

[17] Xiao et al., Kuramoto Oscillatory Phase Encoding (KoPE), ICML 2026, arXiv:2604.07904v2. Primary source.

[18] Trappe, arXiv:2601.04362. Primary source.

[19] CORDIS 101125031, 101092096, 101017098, 871501, 101142702, 101070149. Primary source.

[20] VNN-COMP 2026 results; Kaulen et al., arXiv:2512.19007. Primary source.

[21] Debenedetti et al., Defeating Prompt Injections by Design (CaMeL), arXiv:2503.18813v2 (24.06.2025). Primary source.

[22] Li et al., AgentDyn: Are Your Agent Security Defenses Deployable in Real-World Dynamic Environments?, arXiv:2602.03117v3. Primary source.

[23] MLCommons MLPerf Inference v5.1 (2025); Hugging Face AI Energy Score v2 (2025). Primary source.

[24] Panigrahy & Tyagi, Energy per Successful Goal: Goal-Level Energy Accounting for Agentic AI Systems, arXiv:2605.22883 (2026). Primary source.

[25] Saad-Falcon et al., Intelligence per Watt, arXiv:2511.07885 (version 6, September 2026). Primary source.

Starting components in our ecosystem

These existing tools provide a starting point. The integrated KYMA system remains proposed.

Discuss the research

We welcome scientific collaboration in abstraction, formal checking, planning and resource measurement. Collaboration, doctoral supervision, recruitment and independent assessment require separate agreements; no institution or funded position is committed.

Ways to collaborate · Contact ANULUM · Funding & Grants