Skip to main content

Reflecting on the CommonsDB Final Conference

Highlights from the CommonsDB final conference at EUIPO Alicante, exploring how open standards and federated networks can deliver scalable copyright infrastructure for Europe.

Doug McCarthy

On September 14, 2026, the CommonsDB consortium organized the project's final conference, hosted by the European Union Intellectual Property Office (EUIPO) in Alicante. The event brought together policymakers, technical architects, AI model developers, cultural heritage professionals, and copyright scholars to learn about the outcomes of the CommonsDB pilot and discuss the future of European copyright infrastructure.

Left-to-right: Philippe Rixhon (Valunode), Luna Schumacher (Pictoright), Krishna Sood (Black Forest Labs), and Krzysztof Nichczynski (European Commission – DG CNECT). Photo: Thom Vaughan.

The Infrastructure Challenge: Bridging the Rights Information Gap

Automated systems increasingly make split-second copyright decisions across billions of works—from upload filtering on online platforms to assembling AI training corpora. Yet, because metadata is routinely stripped from digital assets when they are circulated online, reliable rights information remains fragmented or entirely missing. As Paul Keller (Open Future) noted in his opening remarks, tracing CommonsDB back to a 2021 white paper co-authored with Felix Reda on Article 17 of the CDSM Directive, the absence of machine-readable rights metadata is a critical infrastructure problem.

CommonsDB – Verifiable Information Infrastructure That Scales

Over its 18-month run, the CommonsDB pilot has proved that purpose-built, content-derived registries operate reliably at production scale. In the opening session, Doug McCarthy (Open Future), Sebastian Posth (Liccium), and Karin Glasemann (Wikimedia Sverige) highlighted that the registry processed over 6.5 million verified declarations from five data suppliers: consortium partners Europeana Foundation and Wikimedia Sverige, along with the Netherlands Institute for Sound and Vision, the National Gallery of Denmark (SMK), and the Rijksmuseum (Amsterdam).

The pilot validated three core architectural principles:

  • Content-Derived Identification: Using the ISCC standard (ISO 24138), identifiers are generated locally directly from media files ex-post. This keeps rights metadata bound to the asset even when embedded file tags are stripped during distribution.
  • Verifiable European Trust: Declarations are cryptographically signed using eIDAS-compliant qualified certificates and Verifiable Credentials (VCs), providing a verifiable audit trail that links declarations back to authentic declarer identities.
  • Local Peer-to-Peer Synchronization: By separating lightweight registry records from full metadata packages, high-volume users—such as AI developers or online platforms—are able to continuously synchronize registry nodes peer-to-peer using open protocols like Iroh. Posth demonstrated that an entire synchronized local node carrying 6.5 million records requires only regular hardware (such as a Mac mini) to perform zero-latency lookups.

From Centralized Databases to Polycentric Ecosystems

A recurring theme throughout the event was the rejection of centralized, single-source databases in favor of a distributed, federated network of community-governed registries. Carlos Luna García (EUIPO) drew on EUIPO's management of TMView (>140 million trademarks) and GIView to emphasize that European infrastructure should make discovery global without centralizing ownership of the data itself.

This aligns with the technical framework presented by Titusz Pan (ISCC Foundation), who introduced the ISCC Discovery Protocol. Operating as a polycentric system, the protocol uses append-only public logs (hubs) that allow independent registries to advertise holdings without enforcing central gatekeeping. 

Opening the conference's third session, Paul Keller stressed that one of the key insights of the CommonsDB project is that technical infrastructure must be built as close as possible to the constituencies that hold and maintain data—a principle that directly framed the broader debate over how centralized or decentralized European copyright infrastructure should ultimately be.

Krzysztof Nichczynski (European Commission – DG CNECT) presented key findings from the Commission's feasibility study on Text and Data Mining (TDM) opt-outs under Article 4(3) CDSM Directive and Article 53(1)(c) of the AI Act. The Commission recommended a hybrid architecture—combining a central EUIPO resolution layer with distributed onboarding nodes (such as CMOs or national bodies)—and announced plans for an upcoming pilot in cooperation with the EUIPO.

Expanding on ecosystem requirements, Philippe Rixhon (Valunode) shared insights from an EUIPO mapping study identifying hundreds of rights databases across Europe, emphasizing that connected infrastructure must combine fingerprinting and metadata standards with simple, lightweight tools for resource-constrained institutions. Complementing this, Anna Vuopala (Copyright Infrastructure Task Force) presented the CITF's Status Quo and Way Forward report, outlining 76 criteria across legal, technical, and semantic layers to establish a trustworthy open rights data framework.

Solving Practical Needs for GLAMs, Rights Holders, and AI Developers

The conference highlighted how registry infrastructure delivers tangible benefits across diverse stakeholders:

  • Cultural Heritage & Repositories: Karin Glasemann and Hugo Manguinhas (Europeana Foundation) explained how ISCC fingerprinting can help resolve cataloging challenges. While only 25% of Europeana collections contain recognizable persistent identifiers (PIDs), ISCC fingerprints can provide deterministic identifiers across the entire corpus. Glasemann illustrated how CommonsDB surfaces historical licensing variations (such as legacy CC BY labels versus public domain status) and transparently exposes "who declared what when." Additionally, fingerprinting enables automated duplicate detection in Wikimedia Commons upload workflows.
  • Visual Rights Holders: Luna Schumacher (Pictoright) outlined the perspective of visual creators, emphasizing the need for functional opt-out infrastructure today rather than years down the road. She also stressed that voluntary registries must not inadvertently turn into mandatory formalities for creators.
  • AI Developers: Krishna Sood (Black Forest Labs) presented the perspective of frontier visual AI developers, noting that model developers urgently require low-latency, machine-readable, and work-level opt-out signals to distinguish genuine rights holder reservations from domain-level publishers.

Legal Grounding and Next Steps

Addressing the legal framework, João Pedro Quintais (IViR) confirmed that CommonsDB's declarative, non-constitutive architecture carries minimal legal liability exposure under EU copyright law. While voluntary transparency layers fit comfortably within existing mandates, expanding public agencies into authoritative legal status determination would require explicit legislative backing.

The conference concluded with appreciation from Paul Keller and Sebastian Posth for the EUIPO hosts, European Commission officers, CommonsDB team members, and external partners.

Participants at the CommonsDB Final Conference

As European discussions surrounding AI transparency and copyright infrastructure continue, the CommonsDB pilot stands as a working proof-of-concept that open standards, cryptographic trust, and federated networks can deliver a scalable, public-interest copyright ecosystem for Europe.

You can view the complete conference slide deck here.