Contents

GPT-6 Astra has arrived. Reading OpenAI's announcement, the improvements are in research, programming, computer use, and work that spans several steps.

Impressive, apparently. Understood.

But if you work in healthcare, you want something more concrete.

  • So what is it actually useful for in practice?
  • What changes for patients?
  • Can it help with studying, or with making teaching material?

So this time I dug into the medical uses in the form of asking Astra.

Voxel art of Hanageruge and a small white AI robot talking across a workbench about where AI is useful in healthcare

Before you read (please)

This post is based on information public as of 6 September 2026. The conversation is a construction based on my research, not a verbatim quotation of anything Astra said.

Published performance, research on other medical AI, and possible future uses are presented separately. Any medical decision has to be checked by a clinician.

First: did Astra actually get better at medicine?

Me: Are you more dependable than earlier AI when it comes to healthcare?

Astra: Improvements have been reported on medically related evaluations. But a higher score on an evaluation and being able to care for patients safely are not the same thing.

The length-corrected score OpenAI published for "HealthBench Professional" is 60.5% for GPT-5.6 Sol and 63.4% for GPT-6 Astra. On that evaluation, Astra is 2.9 points ahead.

HealthBench Professional (length-corrected score)GPT-5.6 Sol compared with GPT-6 Astra. The gap is 2.9 points
025507510060.5%GPT-5.6 Sol63.4%GPT-6 Astra

This is not a diagnostic accuracy rate. It is an evaluation of situations where a clinician consults an AI: clinical questions, writing and documentation, and medical research. Figures published by OpenAI.

View as a table
ModelScore
GPT-5.6 Sol60.5%
GPT-6 Astra63.4%

This is an evaluation of situations where a clinician consults an AI. It covers not only clinical questions but writing and documentation, and medical research.

So it does not mean "63.4% diagnostic accuracy."

Skipping that and writing "AI finally beats doctors" would be several sizes too big.

What I'm watching is not only answers to medical questions, but the ability to organise researched information and turn it into documents and software. If that grows, the range of uses in healthcare grows with it. What follows are ideas from that angle.

It might cut what you read before a visit and write after it

Me: What looks most useful in day-to-day practice?

Astra: Organising a patient's information, and drafting notes or explanations for a clinician to check.

For instance, asking for this kind of summary before a consultation.

Show me what has changed since the last visit, alongside the original records. Keep anything that needs checking separate from what is already confirmed.

Pulling together tests, medications and information from other departments in a form where you can trace the source. That isn't hypothetical: on 1 September 2026, OpenAI announced an integration for healthcare organisations that lets ChatGPT reference permitted patient information from the Epic electronic health record system.

That said, choosing Astra does not mean it plugs straight into a Japanese hospital's EHR. Separate from what the model can do, you need compatible systems, access rights and safety management in place.

For the notes side, there is actual research, on a different AI product.

A randomised trial published in 2025, with 238 outpatient physicians across 14 specialties. Physicians using the documentation AI Nabla went from 4 minutes 30 seconds per note to 3 minutes 49 seconds. The usual-care arm also improved, from 4 minutes 22 seconds to 4 minutes 4 seconds, but the reduction in the Nabla arm was 9.5% larger, and statistically significant. DAX, evaluated at the same time, showed no significant reduction.

How much did the time per note change?A randomised trial with 238 outpatient physicians across 14 specialties (November 2024 to January 2025, one US medical centre)
  • Before
  • After
01 min2 min3 min4 min5 min4 min 22 s4 min 4 sUsual care4 min 30 s3 min 49 sUsing Nabla

The reduction in the Nabla arm was 9.5% larger than in the usual-care arm, and statistically significant. The DAX arm evaluated alongside it showed no significant reduction. The notes were also reported to contain clinically significant errors occasionally.

View as a table
ArmBeforeAfter
Usual care4 min 22 s4 min 4 s
Using Nabla4 min 30 s3 min 49 s

This is not Astra's record. It also ran at a single medical centre over two months, so the same effect won't necessarily appear everywhere.

More importantly, the study reports that the notes "occasionally" contained clinically significant errors, most often omissions or pronoun mistakes. One mild patient safety event was reported. The research team says explicitly that this needs active physician oversight rather than passive acceptance.

Even so, "the right tool with the right process can reduce paperwork" is a realistic claim, and a much better one than "install AI and everything is solved."

If it were me, I'd want to see whether the time actually dropped once checking and correcting are included, not just whether the draft appeared faster.

Not "summarise this paper" but "make these papers comparable"

Me: What about studying, or reading papers?

Astra: Rather than shortening things to a conclusion, consider using it to organise the differences between studies. A support for lining up populations, comparators, results and limitations.

Something like this.

For these five papers, compare the population, the treatments compared, the observation period, the main results, adverse events and the limitations. Show where the conclusions disagree. If you could not check the full text of a paper, say so.

"Does this treatment work?" gets you one answer.

I would rather have "under what conditions, and how far do we actually know?" laid out.

OpenAI already offers clinician-facing services with cited medical information search and literature review, and has announced connections to public sources such as PubMed and ClinicalTrials.gov for healthcare organisations. Those are mechanisms for reaching information, separate from the Astra model itself.

Which suggests a use that doesn't stop at searching: carry it through to comparison tables, study-group material, checklists.

But don't decide anything from the AI's summary alone. Go back to the original paper for the evidence you'll rely on and any important numbers. I want to keep that step.

Making an explanation fit the person hearing it

Me: What's in it for patients?

Astra: Rewording an explanation a clinician has checked, or helping organise what someone wants to ask at their appointment.

What I picture is less "have the AI diagnose me" and more what happens either side of that.

Take an explanation the clinician has approved and reword it with fewer technical terms. Turn it into a single page a family can share. Before the appointment, organise the history of symptoms and the questions worth asking.

"They explained it at the hospital, but I can't repeat it properly once I'm home."

Having a tool that lets you retrace confirmed information at that moment would help.

In fact, ChatGPT Health for individuals is described as supporting things like tracking changes in test results and organising information before a visit, based on connected health data. Officially it is not intended for diagnosis or treatment; it is positioned as an aid. The official help page I checked on 6 September 2026 says it is for eligible users aged 18 and over in the United States. You cannot assume the same integration is available in Japan.

And this is ChatGPT Health as a service, not a feature unique to Astra.

What I want from this isn't AI removing the conversation with a doctor. It's making that conversation more understandable.

The part I personally find most interesting: building teaching material on the spot

Me: What about anatomy diagrams, or interactive teaching material?

Astra: Building study apps or interactive material on top of trustworthy sources or existing 3D models is a plausible use.

This is where my interest really is.

Imagine material like this. Spin a 3D model of the body, press the part you're curious about and an explanation appears. Look at the hard-to-follow bits from another angle. Same content, but the explanation switches between "for medical students" and "for patients".

Or interview practice with a fictional patient. Instead of handing you the answer, information comes out gradually as you ask. At the end, it walks you through what you failed to ask and where you jumped to a judgement.

None of this means educational effectiveness has been demonstrated with Astra. These are ideas for building and using, extrapolated from the published programming and tool-use abilities.

But as a direction, it's interesting.

From "write me text explaining this" to "build me an experience for understanding this."

And it isn't pure fantasy either. I recently wrote about localising a browser-based 3D anatomy viewer built with Astra. What sits underneath it is body data published by Japanese research institutes.

Of course, I wouldn't call it finished because the screen moves nicely. The structures, the names, the case setup and the explanations all need a specialist to check them. Visual realism is not a substitute for medical correctness.

With that in place, if you can turn a teacher's "this is the bit that never lands" into material, there's a lot of use in it.

So can it be a partner for diagnosis and treatment?

Me: Still, the thing everyone wonders about is diagnosis. What about that?

Astra: As a way for a clinician to check their own thinking — raising alternatives, pointing out missing information — it is plausible. It does not mean an AI's answer alone should decide diagnosis or treatment.

So not "agree with me", but something like this.

Separate the information that supports this judgement from the information that argues against it. List other possibilities worth considering, and what information is still missing.

The research on AI and diagnosis actually contains both encouraging and cautionary results.

Study What was compared Result
Science (2026) A large language model against physicians on clinical reasoning tasks and emergency patient evaluations The model outperformed the comparison physicians' responses. The authors stress the need for prospective trials in real care
JAMA Network Open (2024) 50 physicians split into a group with access to GPT-4 and a group with conventional resources, on case exercises Diagnostic reasoning scores did not improve significantly

Neither is a trial of Astra.

But together they make a point worth sitting with.

An AI producing good answers on its own, and people using AI to deliver better care, have to be confirmed separately.

In the material I checked, I could not find a clinical trial directly showing that using Astra improved outcomes for real patients. Better scores on medical evaluations are a reason for hope, but they are not proof of a treatment effect.

Astra does accept image input, but that alone is not grounds for trusting it with CT or MRI interpretation. Being able to handle images and having verified performance for a specific medical use are different things.

A bit further out, how research gets done may change too

Me: Will it touch research into new treatments and drugs?

Astra: Organising literature, examining hypotheses, writing analysis code, comparing results. Supporting that kind of research process is a plausible direction.

This field isn't pure speculation either.

Separately from Astra, OpenAI has GPT-Rosalind for life-science research, announced for analysing literature and biological data and supporting research in combination with specialist tools. Presenting that as an Astra achievement would be wrong, but the direction — AI moving into the working steps of research — is visible.

What I picture isn't "ask the AI and a new drug appears". It's something that helps a researcher get to the next round of verification.

Organise what is known. Find where results disagree. Consider alternative explanations. Set up the analysis that would test it.

If the time that frees up goes into experiments and thinking, the effect on healthcare over the long run could be large.

If you're starting today, start with work that has no patient data in it

Me: So what should someone try first?

Astra: Work you can test without real patient information: studying from public material, building teaching material from fictional cases.

Here's what I'd try first.

Using only this public document as your source, draft an explanation aimed at patients. Reword technical terms, and do not add anything that isn't in the document. Mark where each important statement comes from, and finish with a list of the points a clinician should check.

That makes it easy to see what the text is based on, and where a human needs to look.

Using real patient information is a different stage with different preparation. Not an individual decision: you need your organisation's approval, a defined purpose, contracts, and rules for storage and access. In Japan, one place to check is the Ministry of Health, Labour and Welfare's "Guidelines for the Safety Management of Medical Information Systems, version 7.0 (June 2026)".

"It won't be used for training" and "it is fine to send this patient's information" are separate questions.

And for patients using it: don't let an AI answer become a reason to delay care you need, or to change medication on your own. Use it as something that connects you to care.

Rather than handing healthcare to AI, I want healthcare to be able to do more

What I kept thinking while researching this: if you judge Astra's arrival only by "will doctors become unnecessary", you miss a lot.

  • Make a patient's course easier to follow
  • Make a clinician's explanation easier to understand
  • Turn papers into material you can compare and think with
  • Turn a teacher's experience into something you can learn by touching

Those are the uses that appeal to me.

The most interesting one: a person with knowledge might now be able to build the thing they kept thinking "if only this existed, it would land."

Not just whether AI knows everything, but whether it can help carry what a person knows to the person who needs it. That's where I think Astra's possibilities in healthcare also lie.

Hope is hope, of course. The effects have to be confirmed for real.

Holding both of those, I think it's fine to be genuinely excited about where AI is going.

Sources

All links checked on 6 September 2026. Availability, features and terms may change.

  1. OpenAI, "GPT-6 Astra: A new generation of intelligence"
  2. OpenAI, "Making ChatGPT better for clinicians" (what HealthBench Professional evaluates, and the clinician-facing services)
  3. OpenAI, "Healthcare organizations can now connect EHR and additional industry data to ChatGPT" (1 September 2026)
  4. Lukac PJ et al., "Ambient AI Scribes in Clinical Practice: A Randomized Trial" NEJM AI. 2025. DOI: 10.1056/AIoa2501000
  5. UCLA Health, "UCLA study finds AI scribes may reduce documentation time and improve physician well-being" (the institution's own summary, 26 November 2025)
  6. OpenAI Help Center, "Health in ChatGPT" (regions, intended use, and that it does not replace care)
  7. OpenAI API documentation, "GPT-6 Astra Model" (intended use, supported input, available tools)
  8. Brodeur PG et al., "Performance of a large language model on the reasoning tasks of a physician" Science. 2026;392:524–527. DOI: 10.1126/science.adz4433
  9. Goh E et al., "Large Language Model Influence on Diagnostic Reasoning: A Randomized Clinical Trial" JAMA Network Open. 2024;7(10):e2440969. DOI: 10.1001/jamanetworkopen.2024.40969
  10. OpenAI, "Introducing new capabilities to GPT-Rosalind" (3 June 2026; a different model, cited as research support)
  11. Ministry of Health, Labour and Welfare, "Guidelines for the Safety Management of Medical Information Systems, version 7.0 (June 2026)" (Japanese)

This post is not advice about any specific treatment or diagnosis. For medical decisions, please consult a clinician.