QFlowLearn: how we built QTI 3 authoring

Open assessment: public evidence review

QTI 3 implementation matrix

LongsightGroup/qti3 implements Question and Test Interoperability (QTI) items and a limited test-sequencing profile through application programming interfaces (APIs). If you evaluate assessment software for your institution, use this matrix to check which capabilities have public test evidence and which need further testing.

By Sam Ottenhoff Published Updated

How to read the matrix

Automated tests show how the repository behaves. They do not show that two independent products can exchange the same content. Certification requires a separate 1EdTech process.

Longsight reviewed version 0.13.1 on October 5, 2026. The links point to the source code used in this review.

What changed in this review

Version 0.13.1 adds combined scoring and restoration scenarios with fixed expected answers. Examples include partial credit for an ordered response, a negative score for an incorrect match, and an adaptive item that remains completed after restoration.

Migration tests now check preserved language, content order, text-entry constraints, and hotspot descriptions. They also check refusals when conversion would discard modal feedback, a response, or a matching restriction. The migration guide explains the acceptance decisions these examples support.

These checks do not establish independent content exchange or accessibility conformance for a complete assessment experience. Certification approval remains pending.

Download the test commands, results, and review limits

Status definitions

Implemented
Public code handles the named feature. This label does not require test evidence.
Tested
Public automated tests cover the claimed behavior.
Demonstrated
A public, repeatable example shows the behavior.
Partial
Only the stated subset has evidence.
Pending
Submitted for external review; awaiting a decision.
Not implemented
The project does not provide the feature.
Not verified
Public evidence is insufficient.
Not applicable
The host application provides the feature.

Evidence matrix

Public qti3 capabilities, evidence, and limitations reviewed October 5, 2026
Area Status Public scope Evidence Last verified Boundary
QTI 3 item XML parsing Tested The core parses item XML into typed models and reports structured diagnostics. Core package and tests This status applies to the repository’s item profile, not every QTI test or results document.
Runtime XSD schema validation Not implemented The qti3 runtime does not run XSD validation. README non-goals Validate XML separately with the official validator or a local schema set pinned to a specific version.
Separate QTI 3 schema checks Tested A separate check validates selected test, rubric, and processing fixtures against official QTI schemas whose file hashes are pinned. Schema validation script This check covers the named fixtures. It does not validate every imported package or establish certification.
Semantic item validation Tested Automated tests cover declarations, response types, references, constraints, and processing diagnostics. Core source Semantic checks do not establish schema validity or destination support.
Standard item interactions Tested The public README lists test fixtures, conformance checks, accessibility metadata, and browser tests for the supported interaction types. Question-type support matrix The support claim covers individual items. Test the surrounding host application separately.
Browser item rendering Tested A native custom-element player renders one item and includes unit and browser test suites. Player package The player is not a complete assessment runner or candidate application.
Response processing and scoring Tested Tests cover response processing, image-area scoring, type checks, numeric overflow, and saved state. Invalid processing produces diagnostics instead of a successful score. Scoring regression tests Host systems own authoritative storage, attempt policy, authorization, and final grade return.
Combined scoring and restoration scenarios Tested Synthetic challenge items combine partial credit, negative scores, numeric boundaries, reused choices, Unicode case rules, overlapping targets, and adaptive completion. Tests compare scores with fixed expected answers and restore saved responses. Challenge questions and expected results These are selected, MIT-licensed test items, not institutional exam data or independent cross-product results.
Item feedback and completed adaptive attempts Tested Tests check that modal feedback follows successful scoring, failed rescoring hides it, and completed adaptive item attempts reject further responses after restoration. Item session source and tests Adaptive item behavior does not provide computer-adaptive test selection. The host still controls review access and attempt policy.
Candidate-safe delivery XML Tested Automated tests cover how the core and command-line interface (CLI) prepare candidate-safe XML and score responses on the server. CLI delivery and scoring commands These APIs do not provide authentication, proctoring, network security, or a full delivery service.
Saved item state Tested Public core and player contracts serialize and restore item responses and state. Repository README The host application must store records, set retention periods, handle concurrent updates, and recover saved state.
Accessibility contracts Partial The repository includes keyboard, focus, accessible-name, validation-message, and proof metadata tests. Accessibility package Automated tests do not prove conformance with the Web Content Accessibility Guidelines (WCAG). You must also test with assistive technology.
Personal Needs and Preferences (PNP) data Tested A package parses, normalizes, validates, and resolves host-provided QTI 3 PNP data. PNP package The host must handle identity, consent, storage, authorization, and service access. Your institution sets accommodation policy.
QTI-shaped authoring output Tested The writer package creates item XML and packages through tested, typed APIs. Writer package A library is not a complete authoring product, review workflow, or content-governance system.
QTI 1.2 and 2.x migration Partial Tests cover specific QTI 1.2 and 2.x conversions, including preserved language, content order, response constraints, and hotspot descriptions. Conversion rejects cases that would lose scoring, modal feedback, response types, or matching-group restrictions. Migrator package Assessment-test structure is not yet preserved; migrated items are returned in a flat review part.
QTI 3 transcoding to standard and product profiles Tested The transcoder has tested export profiles for QTI 1.2, 2.1, 2.2, Canvas, Moodle, Blackboard, and Brightspace. Transcoder profiles and checks Profile tests and schema checks do not prove that every named destination product accepts every generated package.
Package loading and inspection Partial The tools inspect packages and load item references from manifests and assessment-test hierarchies to validate individual items. Command-line interface (CLI) package qti3 provides libraries and tools for assessment delivery platforms. It is not an assessment delivery platform itself.
ZIP and streaming package import limits Tested Core tests cover ZIP parsing and limits on archive resources and streamed input. Browser tests cover stored and deflated ZIP imports, including imports that exceed those limits. ZIP import tests These tests do not cover every archive. Set and test upload limits in your deployment.
Saved-package browser library Tested The playground stores packages locally and preserves their assets. Browser tests cover item submission, scoring, feedback, and reset. Library scoring browser tests Your institution still needs a repository, an authoritative grade record, and backups outside local browser storage.
Finite test sequencing and session replay Tested Core APIs run one linear part with individual submission, flat visible sections, fixed item references, scalar test outcomes, and forward section branches. Both fixed and sequenced tests pass the same execution checks. Test execution acceptance review Execution rejects test time limits, inherited session controls, test feedback and rubrics, nested sections, weights, and random selection. Preserving that data in a package does not mean the runtime can execute it.
Catalog support delivery Tested Tests cover how the core and player resolve catalogs, control requests, and deliver catalog content under host control. 0.9.9 catalog delivery release notes The host still owns catalog authorization, availability, policy, and candidate experience.
Portable Custom Interaction (PCI) host contract Tested The player has tests for parsing PCI launch data and for the APIs a host uses to mount an interaction and manage its state. Portable custom interaction tests The host still owns execution isolation, code trust, accessibility, assets, and product integration.
Assessment delivery platform Not applicable Assessment delivery platforms use qti3 for QTI parsing, validation, rendering, scoring, and supported test sequencing. Library and host responsibilities The platform provides candidate navigation, timing policy, and authoritative attempt storage. These are platform responsibilities, not missing qti3 features.
Shared stimulus delivery Not implemented The repository states that shared stimulus delivery is outside its current item-focused role. README non-goals Test fixtures and package references do not demonstrate shared stimulus delivery.
QTI Results Reporting Not implemented Item outcome state exists, but public evidence does not show a QTI Results Reporting implementation. Repository README Internal outcome serialization is different from implementing the standard reporting role.
Learning Tools Interoperability, roster, identity, and gradebook Not applicable The host provides the learning management system (LMS) interface, rosters, identity, analytics, and gradebook integration. Project scope Test these integrations in the host product and your institutional environment.
Independent cross-product interoperability Not verified The repository has conformance fixtures and external-content tests. We found no current public report that names independent product versions and records complete round-trip results. Conformance package Conformance fixtures are useful implementation evidence, not independent interoperability proof.
1EdTech QTI certification Pending Longsight has submitted its QTI certification application to 1EdTech. Approval is pending. 1EdTech conformance and certification Certification has not yet been granted.

Limits of this guidance

  • This page reviews source code and tests. It does not report independent interoperability testing.
  • Automated accessibility evidence does not establish WCAG conformance for a complete candidate experience.
  • Longsight has submitted its QTI certification application to 1EdTech. Approval is pending.
  • Test product behavior separately in the exact build your institution plans to use.

Missing public test results

  • A cross-product test report with exact versions of independent implementations and representative test content
  • Manual keyboard and assistive-technology results for each high-risk interaction type
  • A migration pilot that names the source and destination builds and records preserved content, scoring results, and rejected items
  • 1EdTech’s decision on the submitted QTI certification application