Google Research Study Reveals How AI Assistance Impacts Professional Expertise
A three-month field study from Google Research and NBER shows that while AI lifts baseline work quality, long-term skill retention differs drastically between junior and senior professionals.

A new three-month field experiment published by the National Bureau of Economic Research and conducted with Google Research reveals that while artificial intelligence tools noticeably improve immediate output, they do not necessarily accelerate long-term skill development for less-experienced workers. The study tracked 133 intellectual property lawyers across eleven firms regularly working with Google to examine how sustained access to an AI patent writing tool—now part of Gemini Notebook—affects both day-to-day productivity and unassisted professional judgment.
Immediate gains in drafting quality
In the randomized controlled experiment, two-thirds of the participating attorneys received access to Google's specialized assistant, while the control group received training on AI concepts but had tool access withheld during the 90-day window. Participants completed drafting tasks based on simulated inventor materials at 10 days and 90 days. Their output was evaluated by independent domain experts across five criteria: enforceability, accuracy, strategic ambiguity, completeness, and clarity.
At 10 days, lawyers with AI access saw drafting scores increase by 0.34 standard deviations (SD), corresponding to a 10-percentile climb compared to the control group. By day 90, this advantage expanded slightly to 0.38 SD, or an 11-percentile boost. Junior lawyers benefited significantly in drafting speed as well, completing the 10-day task 18 minutes faster than the control group's average time of 124 minutes.
However, these quality gains came primarily from moving poor work into average or good territory, rather than expanding the volume of exceptional submissions. Scores moved out of the lowest quintiles into middle bands, leaving the proportion of top-tier work largely unchanged.
Unassisted redlining exposes an expertise gap
To see whether machine assistance built durable capability, researchers added a separate redlining assignment at the 90-day mark. Lawyers had to manually review, annotate, and fix intentional defects and "patent profanity" in a hypothetical filing without any AI tools.
When AI assistance was removed, overall performance remained 0.32 SD higher for the group that had spent three months using the tool. However, this gain was driven almost entirely by senior lawyers, who outperformed controls by 0.45 SD (a 13-percentile jump). Junior lawyers who had used AI showed no clear overall improvement over unassisted peers. Instead, junior scores bifurcated: researchers recorded an increase in very low scores alongside an increase in good scores, with fewer mediocre results and zero gains in excellent results.
Qualitative analysis highlighted distinct working methods between experience levels. Juniors frequently spent excessive time on low-stakes introductory sections, focusing on surface-level word swaps rather than structural legal claims. When identifying serious flaws, junior practitioners often noted the problem in comments without executing the required rewrite—a habit present in both treatment and control junior cohorts that three months of AI assistance failed to correct.
Senior attorneys, on the other hand, treated the AI system as a "logic auditor." Exposure to AI-generated variations prompted experienced lawyers to rethink structural formulations, rebuild core claims from scratch, and justify edits through established legal doctrines, sharpening their existing instincts.
What it means for developers
For engineering teams and technical leaders, the findings highlight an important distinction between augmenting current task throughput and cultivating foundational domain knowledge. In software development—much like legal drafting—less-experienced engineers can use AI code generation to eliminate syntax errors and produce functional code quickly. Yet relying heavily on models to handle implementation may leave junior engineers stuck diagnosing issues without building the tactical instinct required to refactor complex architectures under pressure.
Senior developers, by contrast, possess the mental models needed to use AI effectively as an architectural sparring partner, critiquing model outputs and spotting edge cases that juniors miss. Teams evaluating these dynamics can experiment with different model architectures and workflows to find the right balance between assistance and skill retention. Developers can try top AI models cheaply through one API at https://apixoai.online, making it straightforward to test how different frontier models fit into review pipelines without managing multiple provider accounts.
As Google's researchers note, long-term expertise requires separating the immediate efficiency of generative assistants from an engineer's unassisted capabilities. Building sustainable technical teams will require training workflows that preserve deep hands-on problem solving alongside AI adoption.
Source: Does better work always mean better workers? — Google Research. Written by the Apixo team from that report.
One key for Claude, GPT, GLM, DeepSeek and more. Pay per token with crypto.
Get your API key

