Lee on Synthetic Data

Peter Lee (UC Davis Law) has posted “Synthetic Data and the Future of AI” (110 Cornell Law Review (Forthcoming)) on SSRN. Here is the abstract:

The future of artificial intelligence (AI) is synthetic. Several of the most prominent technical and legal challenges of AI derive from the need to amass huge amounts of real-world data to train machine learning (ML) models. Collecting such real-world data can be highly difficult and can threaten privacy, introduce bias in automated decision making, and infringe copyrights on a massive scale. This Article explores the emergence of a seemingly paradoxical technical creation that can mitigate—though not completely eliminate—these concerns: synthetic data. Increasingly, data scientists are using simulated driving environments, fabricated medical records, fake images, and other forms of synthetic data to train ML models. Artificial data, in other words, is being used to train artificial intelligence. Synthetic data offers a host of technical and legal benefits; it promises to radically decrease the cost of obtaining data, sidestep privacy issues, reduce automated discrimination, and avoid copyright infringement. Alongside such promise, however, synthetic data offers perils as well. Deficiencies in the development and deployment of synthetic data can exacerbate the dangers of AI and cause significant social harm.

In light of the enormous value and importance of synthetic data, this Article sketches the contours of an innovation ecosystem to promote its robust and responsible development. It identifies three objectives that should guide legal and policy measures shaping the creation of synthetic data: provisioning, disclosure, and democratization. Ideally, such an ecosystem should incentivize the generation of high-quality synthetic data, encourage disclosure of both synthetic data and processes for generating it, and promote multiple sources of innovation. This Article then examines a suite of “innovation mechanisms” that can advance these objectives, ranging from open source production to proprietary approaches based on patents, trade secrets, and copyrights. Throughout, it suggests policy and doctrinal reforms to enhance innovation, transparency, and democratic access to synthetic data. Just as AI will have enormous legal implications, law and policy can play a central role in shaping the future of AI.

Kolt on Governing AI Agents

Noam Kolt (University of Toronto) has posted “Governing AI Agents” on SSRN. Here is the abstract:

While language models and generative AI have taken the world by storm, a more transformative technology is already being developed: “AI agents” — AI systems that can autonomously plan and execute complex tasks with only limited human oversight. Companies that pioneered the production of tools for generating synthetic content are now building AI agents that can independently navigate the internet, perform a wide range of online tasks, and increasingly serve as automated personal assistants. The opportunities presented by this new technology are tremendous, as are the associated risks. Fortunately, there exist robust analytic frameworks for confronting many of these challenges, namely the economic theory of principal-agent problems and the common law doctrine of agency relationships. Drawing on these frameworks, this Article makes three contributions. First, it uses agency law and theory to identify and characterize problems arising from AI agents, including issues of information asymmetry, discretionary authority, and loyalty. Second, it illustrates the limitations of conventional solutions to agency problems: incentive design, monitoring, and enforcement might not be effective for governing AI agents that make uninterpretable decisions and operate at unprecedented speed and scale. Third, the Article explores the implications of agency law and theory for designing and regulating AI agents, arguing that new technical and legal infrastructure is needed to support governance principles of inclusivity, visibility, and liability.

Tokson on Government Purchases of Private Data

Matthew Tokson (Utah Law) has posted “Government Purchases of Private Data” (Wake Forest Law Review, Forthcoming) on SSRN. Here is the abstract:

The United States lacks a comprehensive data privacy statute, and most states impose only minimal legal constraints on consumer data collection. This regulatory vacuum has given rise to commercial markets in sensitive private data. In recent years, federal agencies and local police departments have begun to purchase this data from specialized brokers in order to track individuals’ activities over time. Much of this data, collected by cellphone apps and internet servers, is likely constitutionally protected. But government attorneys have mostly concluded that purchasing data is a valid way of bypassing the Constitution’s restrictions.

This Article addresses the increasingly prominent issue of government purchases of private data, and examines broader issues of privacy protection in an era of commercial markets in personal information. The Article questions the widespread assumption that the Fourth Amendment can never apply to commercial purchases. Police officers can generally purchase an item available to the public without constitutional restriction. But a closer examination of data markets demonstrates that sensitive cellphone data is not publicly available or exposed. Rather, the vendors who sell such data do so either exclusively to law enforcement agencies or in large, anonymized chunks to other marketing companies. Because sensitive cellphone data remains functionally private, a government purchase of such data violates the Fourth Amendment.

The Article then challenges the idea that consumers waive their rights in their cellphone data when they use apps or other services. The explanations customers see when an app asks for permission to access their data are often insufficient or misleading, and they typically say nothing about personal data being sold to other parties. Further, penalizing users for disclosing their data to service providers creates harmful incentives and is incompatible with meaningful Fourth Amendment protection in the digital age.

The Article sits at the intersection of consumer privacy and Fourth Amendment law, as poorly regulated markets in personal data and flawed concepts of consumer consent now threaten to erode fundamental constitutional rights. The Article draws broader lessons about the inadequacy of consumer privacy law in the United States. It examines the potential for private surveillance to become government surveillance, via technical and legal interoperability. And it assesses a variety of possible solutions through which legal actors can prevent commercial markets in private data from undermining Fourth Amendment rights.

Re on Artificial Authorship and Judicial Opinions

Richard M. Re (U Virginia Law) has posted “Artificial Authorship and Judicial Opinions” (George Washington Law Review, Forthcoming) on SSRN. Here is the abstract:

Generative AI is already beginning to alter legal practice. If optimistic forecasts prove warranted, how might this technology transform judicial opinions—a genre often viewed as central to the law? This symposium essay attempts to answer that predictive question, which sheds light on present realities. In brief, the provision of opinions will become cheaper and, relatedly, more widely and evenly supplied. Judicial writings will often be zestier, more diverse, and less deliberative. And as the legal system’s economy of persuasive ability is disrupted, courts will engage in a sort of arms race with the public: judges will use artificially enhanced rhetoric to promote their own legitimacy, and the public will become more cynical to avoid being fooled. Paradoxically, a surfeit of persuasive rhetoric could render legal reasoning itself obsolete. In response to these developments, some courts may disallow AI writing tools so that they can continue to claim the legitimacy that flows from authorship. Potential stakes thus include both the fate of legal reason and the future of human participation in the legal system.

Barry on Digital Lawyering: Advocacy in the Age of AI

Patrick Barry (U Michigan Law) has posted “Digital Lawyering: Advocacy in the Age of AI” (Michigan Technology Law Review Forthcoming) on SSRN. Here is the abstract:

All lawyers are now digital lawyers. From Zoom hearings, to e-discovery, to AI-enhanced research and writing, the practice of law increasingly requires the skillful navigation of a wide range of technological tools. It’s no longer enough to be book smart and street smart. More and more, you also have to be byte-smart.

To help future lawyers navigate this transition, I recently created a course at both the University of Michigan Law School and the University of Chicago Law School called “Digital Lawyering: Advocacy in the Age of AI.” The course takes a skill-building approach to artificial intelligence. Which tools are worth using? What questions are worth asking? And how do advocates of all kinds continue to add value to clients—and promote justice—in a world increasingly populated by chatbots, algorithms, and a wide range of other powerful digital products?

This paper collects thoughts from the presentation about the course that I delivered at the “Law and Justice in the Age of AI” symposium organized by the Michigan Technology Law Review on November 18, 2023

Solove on AI and Privacy

Daniel J. Solove (George Washington U Law) has posted “Artificial Intelligence and Privacy” (77 Florida Law Review (forthcoming Jan 2025)) on SSRN. Here is the abstract:

This Article aims to establish a foundational understanding of the intersection between artificial intelligence (AI) and privacy, outlining the current problems AI poses to privacy and suggesting potential directions for the law’s evolution in this area. Thus far, few commentators have explored the overall landscape of how AI and privacy interrelate. This Article seeks to map this territory.

Some commentators question whether privacy law is appropriate for addressing AI. In this Article, I contend that although existing privacy law falls far short of addressing the privacy problems with AI, privacy law properly conceptualized and constituted would go a long way toward addressing them.

Privacy problems emerge with AI’s inputs and outputs. These privacy problems are often not new; they are variations of longstanding privacy problems. But AI remixes existing privacy problems in complex and unique ways. Some problems are blended together in ways that challenge existing regulatory frameworks. In many instances, AI exacerbates existing problems, often threatening to take them to unprecedented levels.

Overall, AI is not an unexpected upheaval for privacy; it is, in many ways, the future that has long been predicted. But AI glaringly exposes the longstanding shortcomings, infirmities, and wrong approaches of existing privacy laws.

Ultimately, whether through patches to old laws or as part of new laws, many issues must be addressed to address the privacy problems that AI is affecting. In this Article, I provide a roadmap to the key issues that the law must tackle and guidance about the approaches that can work and those that will fail.

Choi on AI Malpractice

Bryan H. Choi (Ohio State Law) has posted “AI Malpractice” (73 DePaul Law Review 301 (2024)) on SSRN. Here is the abstract:

Should AI modelers be held to a professional standard of care? Recent scholarship has argued that those who build AI systems owe special duties to the public to promote values such as safety, fairness, transparency, and accountability. Yet, there is little agreement as to what the content of those duties should be. Nor is there a framework for how conflicting views should be resolved as a matter of law.

This Article builds on prior work applying professional malpractice law to conventional software development work, and extends it to AI work. The malpractice doctrine establishes an alternate standard of care—the customary care standard—that substitutes for the ordinary reasonable care standard. That substitution is needed in areas like medicine or law where the service is essential, the risk of harm is severe, and a uniform duty of care cannot be defined. The customary care standard offers a more flexible approach that tolerates a range of professional practices above a minimum expectation of competence. This approach is especially apt for occupations like software development where the science of the field is hotly contested or is rapidly evolving.

Although it is tempting to treat AI liability as a simple extension of software liability, there are key differences. First, AI work has not yet become essential to the social fabric the way software services have. The risk of underproviding AI services is less troublesome than it is for conventional professional services. Second, modern deep-learning AI techniques differ significantly from conventional software development practices, in ways that will likely facilitate greater convergence and uniformity in expert knowledge.

Those distinguishing features suggest that the law of AI liability will chart a different path than the law of software liability. For the immediate term, the interloper status of AI indicates a strict liability approach is most appropriate, given the other factors. In the longer term, as AI work becomes integrated into ordinary society, courts should expect to transition away from strict liability. For aspects that elude expert consensus and require exercise of discretionary judgment, courts should favor the professional malpractice standard. However, if there are broad swaths of AI work where experts can come to agreement on baseline standards, then courts can revert to the default of ordinary reasonable care.

Diamantis on Reasonable AI: A Negligence Standard

Mihailis Diamantis (U Iowa Law) has posted “Reasonable AI: A Negligence Standard” (77 Vand. L. Rev. __ (2025 Forthcoming)) on SSRN. Here is the abstract:

Even as artificial intelligence promises to turbocharge social and economic progress, its human costs are becoming apparent. By design, AI behaves in unexpected ways. That is how it finds unanticipated solutions to complex problems. But unpredictability also means that AI will sometimes harm us. To curtail these harms, scholars and lawmakers have proposed strict regulations for firms developing safe algorithms and strict corporate liability for injuries that nonetheless occur. These rigid “solutions” go too far. They dampen innovation and disadvantage domestic firms in the international technology race.

The law needs a more nuanced framework that balances progress with fairness. Tort law offers a compelling template, but the challenge is to adapt its distinctly human notion of fault to algorithms. Tort law’s central liability standard is negligence, which compares the defendant’s behavior to other “reasonable” people’s behavior. But there is no clear comparison class for AI. Assessing algorithms by reference to people would set too low of a bar—AI can and should outperform reasonable humans on many tasks. Assessing AI instead by reference to itself is often impossible—there are not enough algorithms in many contexts to establish a meaningful baseline.

This Paper offers a novel negligence standard for AI. Rather than compare any given AI to humans or to other algorithms, the law should compare it to both. By this hybrid measure, an algorithm would be deemed negligent if it causes injury more frequently than the combined incident rate for all actors—both human and AI—engaged in the same type of conduct. This negligence standard has three attractive features. First, it offers a baseline even when there are very few comparable algorithms. Second, it incentivizes firms to release all and only algorithms that make us safer overall. Third, the standard evolves over time, demanding more of AI as algorithms improve.