Open Large Language Models: Interdisciplinary Approaches to Embedded Generative AI, Data Protection, and Socio-Legal Responsibility
This research addresses the dominance of centralized large language models (LLMs), where access to advanced latest LLMS is restricted to proprietary platforms (e.g. OpenAI, Claude), and where the datasets used for training remain opaque and largely untraceable. This concentration of control raises significant concerns regarding compliance with European data protection law, intellectual property rights, and the potential use of unlawfully sourced or personal data during model training and fine-tuning.
The project adopts an interdisciplinary approach combining AI engineering, legal analysis, and AI governance frameworks, and proposes the decentralization of large language models through open and distributable architectures. By enabling broader usability beyond original model creators, this approach reduces structural concentration while introducing mechanisms to audit, trace, and partially reverse-engineer training and fine-tuning stages, in order to assess whether illicit or non-compliant data has been incorporated.
The research aims to contribute to ongoing academic and policy debates on AI governance, data protection, and accountability, particularly in the context of the EU Artificial Intelligence Act. It proposes a novel regulatory and licensing framework for open models, based on an open-weights licensing concept, which allows the sharing of model weights and architectures while clarifying the rights and responsibilities of providers, deployers, importers, distributors, and users across the AI value chain.