myofficehours.ai

Issue 6 · 2026-09-18 · 4 headlines · 10 min read

The doing has outrun the checking

This week the models crossed into work, not just answers. A model reconfigured for law ran a competent first pass over nearly the whole published U.S. case-law corpus. Google began handing governments security AI that finds software flaws and applies the fixes itself. The weather shown in Google Search and Maps is now produced by an AI model issuing forecasts every hour, and voice AI that can see your screen and act while you talk arrived in the Google tools your campus already runs. On the same days came the sound of the checking side failing to keep up: only one journal in five has any policy on how reviewers may use AI while submissions flood in, the two largest American school districts pulled generative AI out of their classrooms, and an Australian academic-integrity officer describes pass rates soaring while misconduct cases vanish. The work is being done. The question everywhere at once is who verifies it.

This week’s one thing30 min

Build one class session where students use AI and are graded on their judgement

Read why, then do it ↓

In this issue
  1. 01The counter-current: the largest districts say stop
  2. 02When the marks stop meaning anything
  3. 03Where the evidence actually stands: the medical educators' map
  4. 04One thing to actually try this week

OpenAI's Astra for Law pairs its most capable model with a legal search index built on the Free Law Project's CourtListener collection, which covers more than 99.9 percent of published U.S. precedential case law. On the vendor's own evaluation it passed just over half of 200 legal research questions at its hardest setting, against 39 percent for the same model with ordinary web search. Access is narrow for now, through selected law firms, but the curriculum consequence is already here: when a machine can produce a competent first draft of legal research, finding and checking authorities stops being the first hour of the craft and becomes the whole point of it. Law schools are the first to teach in a field where the entry-level work has been automated, and they will not be the last. Google's releases point the same direction from three sides. WeatherNext 3 trains on live satellite and station observations rather than the output of physics simulations, issues a forecast every hour at five-kilometre resolution, and now supplies the weather in Google Search, Maps and Earth. The underlying grids can be queried in Earth Engine or BigQuery or bulk-downloaded, no model required — which turns a whole family of term projects in agriculture, hydrology and energy studies from build-a-model into query-a-dataset. The Fairwind Program pairs a security-specialised Gemini with tooling that patches the vulnerabilities it finds, aimed first at hospitals, utilities and school districts — the same institutions campuses run. And Gemini 3.8 Live brings voice AI that sees a shared screen and runs tasks in the background to Docs, Gmail and Keep, though only where someone holds a paid Google AI subscription, which quietly turns a campus Google licence into a two-tier system. Against that acceleration, the checking side of the week reads almost as one story told four times. Peer review is triaging harder than it has in years, with desk rejections now outnumbering acceptances two and a half to one and no consensus on what reviewers may do with AI. Two of the three largest American school districts have suspended classroom AI outright. An Australian integrity officer reports a year with essentially no plagiarism cases, collapsing fail rates, and no structural check that the marks mean anything. And the best-controlled study of AI simulated patients in medical education measured a rise in student confidence, with competence still unmeasured. Four fields, one shape: the producing is getting cheaper faster than the verifying is getting better.

01

The counter-current: the largest districts say stop

New York City Public Schools, the largest district in the country, announced a one-year moratorium on student use of generative AI from early childhood through eighth grade, affecting roughly 600,000 students, and is switching off the AI components of more than 38 previously approved programs that do not meet its new safety standards. Los Angeles Unified, the second largest, blocked student access to generative AI on district-issued devices, announced to its own board by the chief academic officer. Miami-Dade, the third largest, went the other way and signed a partnership with Google Gemini, and in Chicago more than half the school-board candidates on the autumn ballot back a moratorium. Ray Schroeder, who gathered this in Inside Higher Ed, thinks the bans are a mistake and that higher education's duty is to train the districts through continuing education. Whatever the right answer, instructors of first-year students will meet it sitting in the room: this year's intake arrives with school AI experience that ranges from daily tutor use to an explicit district ban, depending on geography, and the older equity gap in teacher training is hardening into who has used these tools at all.

02

When the marks stop meaning anything

Kane Murdoch, who runs academic-misconduct cases for a living, published the most uncomfortable paragraph of the week: a colleague who serves as an integrity officer at an Australian university told him the institution has had essentially no plagiarism cases in a year while pass rates skyrocketed, and he traces it to the structure of assessment rather than to cheating detection. Where one supervised task sits alongside tasks where AI is allowed, and the supervised task is not a hurdle that must itself be passed, a student can score 35 out of 50 on the open tasks, 15 in person, and still pass the subject. He ties the standard to section 1.4.4 of Australia's Higher Education Standards Framework, which requires that completed students have demonstrated the learning outcomes — which makes hurdle status a concrete, discussable policy lever rather than a vague worry. The journals are discovering the same squeeze from the other end: submissions rose about 68 percent between 2018 and 2025, reviewer acceptance rates fell every year to 22 percent, and while just over half of reviewers report using AI for review tasks, only about one journal in five has a policy saying what is allowed.

03

Where the evidence actually stands: the medical educators' map

Andrew O'Malley's This Week in MedEd did the reading most assessment committees have not, pulling four peer-reviewed 2026 studies into one map of AI in the simulated patient encounter. The pattern that emerges is consistent: AI handles the observable, repeatable and countable parts of clinical teaching and fails at the relational parts. A chatbot playing the patient raised students' self-rated communication competence, but could not portray an ambivalent, resistant patient — the clearest example yet of AI sycophancy in medical education — and awarded high marks to 85 percent of students with feedback quality that did not correlate with learning gains at all. Automated marking of objective structured clinical exams agreed worst on the communication-dependent elements those exams exist to test, and the reviewers' governance guidance is specific: augmentation only, local validation before adoption, bias audits at fixed intervals, a documented appeal pathway wherever AI-influenced scores affect progression, and no fully autonomous AI scoring in high-stakes decisions. That list is not a medical-education rule; it is a template any faculty can carry to their own assessment committee.

04

One thing to actually try this week

Build one class session where students use AI and are graded on their judgement. Give them an AI-written first pass — a literature summary, a legal memo, a data analysis — and make the graded work the catching: which citations are real, which numbers check out, which confident sentence is wrong. Ethan Mollick's argument this week is that the models can now do weeks of competent work when properly guided, and that what they cannot supply is deep knowledge, wide knowledge, taste and agency — which is exactly what marking the verification measures. The AI literacy task on the dashboard has a free action that walks through building the session, with the prompt and the marking guidance already written.

Do this now30 min

Build one class session where students use AI and are graded on their judgement

A single session plan with an activity, a prompt for students, and a short rubric that assesses the quality of their evaluation rather than the quality of the output.

Most AI policy tells students what they may not do. Almost none of it teaches them to do it well. A survey of 45,398 students and staff across 35 countries found only 29 per cent of students think their instructors can guide them, and that gap is closed one session at a time.

The steps

  1. 01Pick one assignment your students already do, and one point in it where they would be tempted to use a model.
  2. 02Have the model produce a deliberately plausible but flawed answer at that point. Keep it.
  3. 03Run the prompt below to build the session around catching the flaw.
  4. 04Grade the critique, not the output. Say so in the rubric, in writing.

The prompt

I teach [SUBJECT] to [LEVEL] students. Here is an assignment they do: [PASTE].

Produce a plausible but subtly flawed AI answer to it. The flaw must be one a competent student in this field could catch with what I have taught them, and it must not be a factual howler.

Then give me:
1. A thirty-minute session plan in which students find the flaw themselves rather than being told it.
2. Three questions I ask when nobody spots it.
3. A four-line rubric that assesses the quality of their reasoning about the answer, not the answer.

Name the flaw at the end, under a heading, so I can read it after I have tried to spot it myself.

Costs nothing

The route uses openly licensed lesson plans from two university libraries and works entirely on a free account.

Before you do

Do not run the session with an example you have not checked yourself. If you cannot find the flaw, your students will not, and the lesson becomes a demonstration that the model is reliable. Bear in mind too that OpenAI has served ads to logged-in adult users on the free and Go tiers of ChatGPT since February 2026 in the United States, and since 11 August 2026 in the United Kingdom, Mexico, Brazil, Japan and South Korea, matched on the topic of the conversation and on past chats. Pro, Business, Enterprise and Education accounts stay ad-free. If you ask a class to run prompts on personal accounts, tell them that is what is happening rather than letting them find out.

Leans on:

Everything else in Teaching AI literacy →

The expert read

In The Overhang, published the day this issue was written, Mollick argues that the argument about future models is obscuring the ones already here: GPT-6 Astra and Claude Fable 5.1 can reliably do weeks of human work when properly guided, and he demonstrates it by having one rebuild the 1977 text adventure Zork as a playable 3D game and another reconstruct Umberto Eco's five-thousand-book library from photographs and catalogues, shelf by shelf. His name for the gap is the overhang: capability that exists faster than habits, systems and institutions absorb it. For universities the implication cuts both ways — the scarce inputs are now the ones a degree exists to build, deep knowledge, wide knowledge, taste and agency, and the pressure his own examples create is the one the rest of this issue documents: our very human processes for checking work move far more slowly than the work now arrives.

Ethan MollickRead the piece ↗Professor at the Wharton School, University of PennsylvaniaEnthusiastweekly

Where the experts actually disagree

With more than 80 percent of students already using AI for schoolwork, should schools block it in the classroom, and what should universities do about the students that produces?

Ray SchroederEnthusiast

The bans are a mistake. Cutting students off from the tools their future workplaces use protects neither critical thinking nor safety, and higher education should respond not with more prohibition but by training K-12 boards, administrators and teachers through continuing education so districts can build workable policy of their own.

Read it ↗
Zohran Mamdani and Kamar Samuels, New York City Public SchoolsSkeptic

A one-year moratorium for the youngest students is the responsible path. The district is discontinuing the AI components of more than 38 approved programs that do not meet new safety and oversight standards, citing the protection of critical thinking and of student-teacher relationships, and the moratorium buys time to set those standards properly.

Read it ↗

Where it lands

The two positions share more than the headlines suggest: both treat the status quo of unsupervised student AI use as unacceptable. They differ on whether the response is control or preparation, and the decision that actually faces a university this term is narrower than either — what to assume about the first-year intake, whose school AI experience now ranges from daily tutor use to an explicit district ban depending on where they went to school. That assumption has to be made course by course, and it is worth making explicitly rather than by accident.

See all 75 voices and where they disagree →

What we said last week, and what happened

The claims this briefing has already made, and how they have held up since.

  • Issue 5 · 10 September 2026

    We said: Issue 5 closed on the question the assessment field could not yet answer: given that AI can now solve almost any written assignment and cannot be trusted to grade one, what is the right response in a university course?

    The first worked answers arrived this week from three directions, and none of them is detection. Kane Murdoch's account of vanished misconduct cases and soaring pass rates points at assessment governance: make the supervised task a hurdle, or the degree cannot certify what it claims. The medical-education reviewers point at controlled augmentation: local validation, bias audits, a documented appeal pathway, and no autonomous AI scoring where progression is at stake. And the peer-review numbers show the same squeeze reaching journals, where two and a half desk rejections now accompany every acceptance and only one journal in five has said what reviewers may do with AI. The question from last issue has moved from whether to respond to which lever each institution pulls first.

What just became possible

Not things to do this week. Things that can now be done at all.

  • Already shipping

    Hourly global weather forecasts at five-kilometre resolution became a public dataset anyone can query, when Google's WeatherNext 3 began powering Search, Maps and Earth and opened the grids in Earth Engine, BigQuery and bulk download.

    A term project that needs real weather — crop stress, solar-farm output, flood exposure — can now be built by querying a public dataset instead of running a forecast model.

    What to do now The grids are bulk-downloadable from Google Cloud Storage today and Earth Engine is free for research use; the data task on the dashboard has the free starting point for working with a dataset of this shape. Google DeepMind blog

The week in full

The past eight days, weighted by what each story asks of you.

  • quiet·This Week in MedEd (Andrew O'Malley PhD)·2026-09-14Act

    Can AI play the patient?

    Andrew O'Malley's This Week in MedEd newsletter (14 September 2026) reviews four peer-reviewed 2026 journal articles, not preprints, that together map how AI is entering the simulated patient encounter in undergraduate medical education: writing the case, playing the patient, marking the station, and running the debrief. A scoping review by Thind and colleagues (Frontiers in Digital Health) screened 1,130 records and found 39 implemented AI innovations in clinical skills curricula between January 2022 and January 2026; 19 were virtual patients powered by large language models (the technology behind chatbots like ChatGPT), 25 of the 39 appeared in 2025 alone, and while feasibility and acceptability were consistently reported, long-term retention and transfer to real patients were almost never measured. A quasi-experimental study by Müller and colleagues (JMIR Medical Education) gave 162 medical students at Charité one conversation (median 22 minutes) with a GPT-4o chatbot playing a patient; self-rated communication competence rose from 5.89 to 6.83 on a 0–10 scale (a moderate effect, Cohen's d=0.58), with significant gains in three of four scenarios but none for motivational interviewing, because the chatbot could not portray an ambivalent, resistant patient — which O'Malley calls the clearest practical example of AI sycophancy, a chatbot's built-in urge to be agreeable, in medical education. That study had no control group and self-reported outcomes, so it measures confidence rather than competence; the chatbot also awarded 85% to 92 of the 162 students, and its feedback quality did not correlate with learning gains at all. A second scoping review, by León-Ariza and colleagues (Medical Teacher), covered 22 studies of AI marking OSCEs, the staged practical exams where students rotate through timed stations with actor patients: automated video scoring produced higher and more uniform marks than human examiners (for example, knot tying 16.07 versus 10.44), agreed best on visually observable tasks and worst on communication-dependent elements (transcript-based agreement 26% to 83%), and while 70% of faculty rated AI-simulated patients good or excellent, 85% flagged the missing non-verbal communication and the impossibility of being physically examined; the authors conclude AI currently augments rather than transforms OSCE assessment, and recommend local validation before adoption, bias audits at fixed intervals, a documented appeal pathway where AI-supported scores affect progression, and no fully autonomous AI scoring in high-stakes decisions. A narrative review by Khan and colleagues (Advances in Medical Education and Practice) found a decade of AI-supported debriefing research is still feasibility studies and prototypes with no validated metrics, and endorses a hybrid model in which AI prepares transcripts and event highlights while a human facilitator keeps the debrief; it also warns that speech-to-text systems mis-transcribe non-native accents. O'Malley's synthesis: AI handles the observable, repeatable, countable parts of simulation and fails at the relational parts, so use it to multiply low-stakes practice repetitions and keep the actor, the examiner and the debrief for the encounter that counts. He discloses his own AI patient simulator, SimPatient, and a related 2025 commentary in Simulation in Healthcare.

    Anyone running clinical skills teaching gets the first consolidated map of where AI patients help — cheap, repeatable, low-stakes practice — and where they fail: ambivalence, non-verbal cues, and anything requiring physical examination. Assessment leads get concrete governance guidance: AI marking agrees worst on the communication elements OSCEs exist to test, and the reviewed evidence supports augmentation only, with local validation, bias audits and an appeal pathway before any AI-influenced score affects progression. Because the flagship study had no control group and self-reported outcomes, the honest reading is that AI patients raise confidence; whether they raise competence is still unmeasured.

New in the library

New

7 resources added since issue 5

Teaching & course design

Data & analysis

Images & figures

Grants & funding

Ethics & compliance

Catching up

These ran earlier than this week and still matter — 4 stories worth a minute of scanning.

  • OpenAI·9 days agoOwn announcement

    Introducing Astra for Law

    AI now does first-pass U.S. legal research over nearly all case law; law faculty must teach students to verify every cited authority.

  • Google DeepMind blog·15 days ago

    Introducing WeatherNext 3, our most advanced and accurate global weather AI model

    Weather in Google Search and Maps is now AI-generated; researchers can query hourly 5-km forecast data in Earth Engine or BigQuery.

  • Google DeepMind·16 days agoOwn announcement

    Proactive cyber defense for governments and enterprises

    Google is handing governments AI that finds and fixes security flaws autonomously; university networks and security teaching will face the same shift.

  • Guerilla Warfare (Kane Murdoch)·228 days ago

    On Bondage

    Soaring pass rates may be AI, not learning: where in-person assessments aren't hurdles, students can pass without demonstrating anything.

Where to go next

The brief is the week. These three are everything behind it.