Title
Author
DOI
Article Type
Special Issue
Volume
Issue
1Department of Anesthesiology and Reanimation, SBU Adana City Training and Research Hospital, 010650 Adana, Türkiye
*Corresponding Author(s):ugur.citilcioglu@saglik.gov.tr (Uğur Serkan Çitilcioğlu)
| History | Submitted: 24 January 2025 | Accepted: 15 May 2025 | Published: 08 January 2026 |
| Copyright: | ©2026 The Author(s). Published by MRE Press. |

Background: This study aimed to evaluate the performance of ChatGPT in interpreting arterial blood gas results and compare it to the performance of physicians with varying levels of experience. Methods: In this comparative and cross-sectional study, 30 selected clinical cases encompassing simple and mixed acid-base disorders, respiratory abnormalities, and electrolyte disturbances, were analyzed by ChatGPT and 45 anesthesiology physicians who were divided into three groups: specialists, experienced residents, and inexperienced residents. Participants assessed arterial blood gases across five domains, including acid-base diagnosis, respiratory evaluation, fluid and electrolyte disturbance, other abnormalities, and treatment planning. Both ChatGPT and participant responses were scored by two experienced anesthesiology academicians based on a 5-point Likert scale. The overall score was calculated as the sum of the scores from the five individual subdomains. Results: The overall scores of ChatGPT and the physicians were similar (21.75, 21.97). ChatGPT demonstrated high accuracy in diagnosing primary acid-base disorders and providing treatment recommendations, though slight variability was observed in mixed disorders. A strong correlation was observed between ChatGPT’s scores and those of physicians (r = 0.912, p < 0.001). Conclusions: ChatGPT demonstrated a performance comparable to that of physicians, suggesting that it could assist in the decision-making processes of healthcare professionals and contribute to workload reduction.
Cite this article
Uğur Serkan Çitilcioğlu, Barış Arslan, Hatice Kaya Özdoğan, Hakan Yalım, Ümit Kara. Is ChatGPT able to interpret arterial blood gas analysis?A comparative cross-sectional study. Signa Vitae. 2026; 22(1): 58-64. doi: 10.22514/sv.2025.138
Arterial blood gas (ABG) analysis is a cornerstone in critical care and anesthesia practice, offering vital information about a patient’s respiratory and metabolic status. ABG interpretation is crucial for monitoring ventilation, oxygenation, and acid-base balance during surgeries, particularly in high-risk patients or complex procedures [1]. However, interpreting ABG results can be challenging, even for experienced clinicians, due to the sheer volume of data and the need for rapid and accurate decision-making in a dynamic clinical environment [2]. One of the significant challenges in ABG interpretation lies in synthesizing multiple parameters, such as pH, pCO2 (partial pressure of carbon dioxide), pO2 (partial pressure of oxygen), HCO3− (bicarbonate), electrolytes, glucose, and lactate, while simultaneously considering the patient’s clinical context [3]. Errors or oversights can occur, particularly under time pressure, leading to delayed or suboptimal treatment decisions [4]. Moreover, less experienced clinicians may struggle to identify subtle abnormalities or formulate appropriate treatment recommendations, further complicating patient management [5].
Artificial intelligence (AI) has emerged as a transformative tool across various medical specialties, offering new opportunities to address such challenges. Among AI systems, ChatGPT, developed by OpenAI, has generated considerable excitement due to its ability to process and synthesize complex information efficiently. ChatGPT is a large language model trained using reinforcement learning from human feedback (RLHF) (OpenAI, https://openai.com/blog/chatgpt), allowing it to interact with users and provide outputs based on prompts. Since its launch in November 2022, ChatGPT has been explored for its potential applications in medicine, including clinical and laboratory diagnostics, patient management, and medical education [6, 7, 8, 9].
AI has demonstrated success in specialties such as radiology and pathology, where large datasets and structured data drive decision-making [10, 11]. However, early studies have highlighted the limitations of AI systems, particularly concerns regarding performance variability and bias in clinical diagnostics [12, 13, 14, 15, 16]. According to the available literature, its potential in dynamic, context-heavy tasks such as ABG interpretation remains unexplored. Due to the complexity of ABG analysis, especially in cases involving mixed acid-base disorders or coexisting respiratory and metabolic abnormalities, AI systems may play a significant role in reducing cognitive workload and improving accuracy [17]. By systematically interpreting ABG results and providing evidence-based treatment recommendations, ChatGPT could assist clinicians in rapid decision-making, enhancing both patient safety and clinical efficiency [4]. Furthermore, ChatGPT offers potential as an educational tool, helping residents and junior clinicians strengthen their understanding of ABG interpretation through case-based learning and interactive feedback [8, 9].
This study aims to evaluate ChatGPT’s performance in interpreting ABG results by comparing it with a group of physicians. Additionally, it explores ChatGPT’s potential as a valuable decision support tool for enhancing patient safety and clinical efficiency in the fields of anesthesia and intensive care.
This comparative and cross-sectional study was conducted at SBU Adana City Education and Research Hospital between 02 January and 20 April 2025, after obtaining ethical approval. The study utilized blood gas analysis results and the clinical medical histories of patients who underwent surgical procedures or were admitted to the intensive care units of the hospital during this period. A total of 160 ABG analyses and patient histories were reviewed, from which 30 representative cases, illustrating both primary and mixed acid-base disorders, were selected during a session organized by the authors.
Inclusion criteria for case selection were as follows:
● Adult patients (≥18 years) who underwent arterial blood gas (ABG) analysis during the study period,
● Availability of complete clinical documentation, including demographic data, clinical history, and ABG values,
● Representation of distinct acid-base disorders, including both primary and mixed disturbances.
Exclusion criteria included:
● Pediatric patients (<18 years of age),
● Incomplete or missing ABG data,
● Lack of sufficient clinical information (e.g., undocumented comorbidities, medications, or presenting complaints).
For each case, ABG results were presented to ChatGPT (GPT-4.0, OpenAI, San Francisco, CA, USA) in the form of photographic images taken directly from printed outputs of the hospital’s blood gas analyzer. These images reflected real-world ABG printouts, preserving original formatting and values. Additionally, each case included a separate photograph of a handwritten note summarizing the patient’s demographic data and brief clinical history. The model was tasked with diagnosing abnormalities and providing treatment recommendations based on five predefined domains (Table 1). To ensure unbiased outcomes, all responses were documented following the model’s initial inquiry, as ChatGPT retains contextual memory within a single session.
| Diagnosis/Assessment | Description |
| Acid-Base Disorder Diagnosis | Determining the type of disorder (respiratory, metabolic, mixed), whether it is compensated or not, and whether it involves acidosis or alkalosis. |
| Fluid and Electrolyte Disorder | Identifying abnormalities in electrolytes or fluid balance. |
| Respiratory Parameters Evaluation | Assessing for hypoxia or hypercapnia. |
| Other Anomalies | Highlighting any abnormalities such as elevated lactate, carboxyhemoglobin (CO), methemoglobin, or glucose levels. |
| Treatment Recommendations | Proposing treatment strategies based on the identified abnormalities. |
The cases were presented to physicians in the same format. Participants were instructed to analyze the cases and provide answers under the same five domains. The study included 45 physicians with varying levels of experience in intensive care and anesthesia (Table 2).
| Group | Participants | Description |
| Specialist Group | 15 | Physicians with at least 5 years of experience in anesthesia and intensive care |
| Experienced Resident Group | 15 | Residents with 2 or more years of training in anesthesia and intensive care |
| Inexperienced Resident Group | 15 | Residents with less than 2 years of training |
A standardized 5-point Likert scale (Table 3) was used to assess each response, as commonly applied in prior studies evaluating AI performance in clinical decision-making [18, 19, 20].
| Score | Description |
| 1 | Completely incorrect/Irrelevant response |
| 2 | Mostly incorrect with minimal relevant information |
| 3 | Partially correct but incomplete or vague |
| 4 | Mostly correct with minor omissions or inaccuracies |
| 5 | Completely correct and clinically appropriate |
The assessment of acid-base disorders was performed using a standardized three-step approach: (1) primary disorder identification (pH, PCO2, HCO3− thresholds), (2) compensation analysis (Winter’s formula, respiratory compensation rules), and (3) anion gap evaluation (calculation and delta ratio interpretation), utilizing established physiological ranges and widely accepted clinical equations [21].
The following diagnostic thresholds were applied for other parameters based on established clinical guidelines: hyponatremia was defined as serum sodium <135 mmol/L and hypernatremia as >145 mmol/L; hypokalemia as potassium <3.5 mmol/L and hyperkalemia as >5.0 mmol/L; hypoglycemia as glucose <70 mg/dL and hyperglycemia as >140 mg/dL; hypoxia as arterial partial oxygen pressure (PaO2) <80 mmHg; hyperlactatemia as lactate >2.0 mmol/L; and clinically significant methemoglobinemia as >1.5% of total hemoglobin [22, 23, 24, 25].
The treatment of acid-base disorders was evaluated according to four fundamental principles: (1) targeted treatment of the underlying etiology, (2) correction of coexisting blood sugar, fluid and electrolyte abnormalities, and (3) implementation of appropriate respiratory interventions when necessary (including mechanical ventilation for severe respiratory acidosis or supplemental oxygen for hypoxemia), (4) recognition of hemoglobinopathies (methemoglobinemia) [26, 27, 28].
Two experienced anesthesiology academicians (BA, with 15 years of experience, and HKO, with 25 years of experience) evaluated the responses provided by ChatGPT and the physicians using the 5-point Likert scale. To ensure objectivity, all evaluations were conducted independently and blinded to participant identity. Also, responses generated by ChatGPT were transcribed into standardized evaluation forms without any indication of their origin, ensuring that evaluators were unaware of whether a response came from a human participant or the AI model. Each category was scored by the two raters based on the clinical accuracy and appropriateness of the response. To assess the inter-rater reliability between the two independent raters, Cohen’s kappa (κ) statistic was calculated. The overall kappa value was found to be 0.82, indicating almost perfect agreement between the raters.
In total, 1380 individual assessments were performed, covering the interpretation of 30 distinct ABG cases by 45 physician participants and ChatGPT. Each case was evaluated across five predefined clinical domains. In cases of scoring discrepancies, the average score was considered final.
Continuous variables are presented as mean ± standard deviation (SD) or median (interquartile range (IQR)), while categorical variables are expressed as numbers (proportions). Normality of continuous variables was assessed using the Shapiro-Wilk test. Comparisons of continuous variables between groups were performed by Kruskal-Wallis and Mann-Whitney U-tests. For categorical data, comparisons between groups were made using the chi-square test or Fisher’s exact test. Correlation analysis between scores was conducted using Spearman’s rank correlation coefficient. All statistical analyses were performed using the Statistical Package for Social Sciences (SPSS; version 18.0; IBM Corp., Chicago, IL, USA), and a p-value of < 0.05 was considered statistically significant.
The median overall score for all participants was 21.97 (IQR: 20.9–23), indicating a generally high performance across groups. No statistically significant differences were observed among the three participant groups regarding total score or individual parameter scores (p > 0.05). ChatGPT’s total score was 21.75 (IQR: 20.7–22.2), closely aligning with the median score of participants (Fig. 1).

Fig. 1.Overall performance of the participant groups and ChatGPT.
The median scores for acid-base disorder diagnosis were 4.3 for specialists, 4.5 for experienced residents, and 4.3 for inexperienced residents. ChatGPT received a median score of 4.3, indicating a performance that was comparable to that of the human participants.
Physicians and ChatGPT both demonstrated high accuracy in diagnosing primary acid-base disorders. For mixed acid-base disorders, participants achieved a mean score of 3.9 ± 0.8 out of 5, while ChatGPT scored 4 ± 1.2. Despite slightly higher variability, ChatGPT’s performance was consistent with the participant groups (Fig. 2).

Fig. 2.Performance in simple and mixed acid-base disorders.
The accuracy of respiratory issue diagnosis, electrolyte disturbance identification, treatment planning, and detection of other anomalies was consistent across all participant groups (p > 0.05). Detailed scores for these categories are summarized in Table 4. ChatGPT demonstrated similar performance to the participants across all parameters (Table 5).
| Variable | Specialist Group (Median (IQR)) | Experienced Resident Group (Median (IQR)) | Inexperienced Resident Group (Median (IQR)) | p-value |
| Diagnosis of acid-base disorder | 4.3 (4.1–4.7) | 4.5 (4.4–4.6) | 4.3 (4.1–4.4) | 0.056 |
| Diagnosis of respiratory issue | 4.5 (4.2–4.7) | 4.6 (4.4–4.9) | 4.3 (3.9–4.5) | 0.062 |
| Diagnosis of electrolyte disorder | 4.4 (4.2–4.6) | 4.5 (4.4–4.7) | 4.3 (4.0–4.8) | 0.265 |
| Treatment plan | 4.3 (4.0–4.7) | 4.4 (4.3–4.6) | 4.2 (3.7–4.3) | 0.053 |
| Other issue | 4.5 (4.2–4.6) | 4.5 (4.4–4.8) | 4.3 (4.0–4.6) | 0.054 |
| Overall score | 22.0 (20.9–23.0) | 22.5 (21.7–23.0) | 21.4 (20.4–22.1) | 0.067 |
| ● Data are presented as medians with interquartile ranges (IQR). | ||||
| ● p-values were calculated to assess statistical differences between groups. |
| Variable | Participants | ChatGPT | p-value |
| Diagnosis of acid-base disorder | 4.37 (4.10–4.70) | 4.30 (4.10–4.70) | 0.243 |
| Diagnosis of respiratory issue | 4.47 (4.10–4.70) | 4.34 (4.10–4.70) | 0.194 |
| Diagnosis of electrolyte disorder | 4.40 (4.10–4.70) | 4.45 (4.10–4.70) | 0.354 |
| Treatment plan | 4.30 (4.10–4.70) | 4.21 (4.10–4.70) | 0.390 |
| Other issue | 4.43 (4.10–4.70) | 4.45 (4.10–4.70) | 0.574 |
| Overall score | 21.97 (20.90–23.00) | 21.75 (20.70–22.20) | 0.632 |
| ● Data are presented as medians with interquartile ranges (IQR). | |||
| ● p-values were calculated to assess statistical differences between groups. |
No statistically significant differences were observed among the specialist, experienced residents, and inexperienced resident groups across any subcategories (p > 0.05) (Table 4). There were no statistically significant differences between ChatGPT and the participants across the five subdomains (Table 5). A strong correlation was observed between ChatGPT’s scores and those of physicians (r = 0.912, p < 0.001).
In this comparative cross-sectional study, we found that ChatGPT demonstrated performance comparable to that of physicians with varying levels of experience in interpreting ABG results.
ABG analysis is a challenging skill in anesthesia and intensive care medicine, where timely and accurate interpretation of acid-base, respiratory, and metabolic abnormalities is essential. This challenge becomes even more apparent, particularly in scenarios where quick decision-making is crucial [29]. Therefore, artificial intelligence, with its ability to perceive details and perform rapid analysis, can significantly ease the workload of clinicians. In this study, both participants and ChatGPT demonstrated high accuracy in diagnosing simple acid-base disorders. For mixed acid-base disorders, which are inherently more complex, participants achieved slightly lower scores compared to primary disorders. ChatGPT’s performance in these cases was also variable, but within the range of human participants. For instance, one representative case involving a patient who had undergone resuscitation following an anaphylactic shock. The arterial blood gas revealed a mixed acid-base disorder characterized by respiratory acidosis (pH: 7.230, pCO2: 59.6 mmHg) and borderline metabolic acidosis (HCO3−: 20.1 mmol/L, Base Excess: −2.9), along with borderline lactate elevation and hypoxemia (Lactate: 4 mmol/L, pO2: 69 mmHg). Clinically, this profile was compatible with the patient’s post-resuscitation status and tissue hypoperfusion. However, despite receiving the full clinical context, ChatGPT categorized the case solely as a primary respiratory acidosis and failed to recognize the coexisting metabolic component and the clinical implications of rising lactate levels. This may indicate a limitation in the model’s ability to integrate complex physiological patterns, particularly in the absence of explicit prompting. In a real-world scenario, such a misclassification could obscure the diagnosis of an ongoing shock state, potentially leading to delays in hemodynamic support or further diagnostic work-up. This example underscores the importance of human oversight in interpreting AI outputs, particularly in cases involving multiple, interrelated pathophysiological processes. Our findings are consistent with previous research. Earlier studies have demonstrated that ChatGPT performs quite well in solving problems that are well-structured and less detailed, but its success rate can fluctuate in scenarios involving rare conditions and complex details [30, 31].
In our study we find strong correlation between ChatGPT’s scores and participant performance, which highlights its potential as a decision support tool for clinicians. Our study also highlights that, in its current state, ChatGPT exhibits slightly lower overall performance compared to human participants, although this difference is not statistically significant. This finding underscores the importance of having its recommendations reviewed by a clinician in clinical decision-making processes to ensure patient safety. A previous study highlighted the potential of using a hybrid system where clinicians leverage AI support in clinical judgment processes. When this application is integrated into clinician’s workflows, it has been observed to significantly enhance their ability to manage patients and make balanced and accurate decisions [32].
Recent advancements in artificial intelligence have demonstrated significant potential in automating repetitive and data-intensive tasks, enabling clinicians to focus on more complex and critical decision-making processes [33]. For example, Wong et al. [31] (2021) validated an AI-based sepsis prediction model that performed comparably to expert clinicians, showcasing the potential of AI in critical care decision-making. Opportunities presented by AI, such as improving medical imaging diagnostics and reducing human errors, are particularly noteworthy [34, 35, 36]. As healthcare continues to evolve, integrating AI frameworks will be essential for addressing the challenges of modern medicine while maintaining the critical human connection in patient care [37]. When integrated with electronic healthcare systems (EHS), ChatGPT can facilitate rapid data analysis and provide instant feedback in point-of-care (POC) applications [38]. In fields such as anesthesia and critical care, it has the potential to reduce cognitive load and minimize errors by offering systematic arterial blood gas (ABG) interpretations and evidence-based treatment recommendations.
Artificial intelligence has certain limitations, ChatGPT’s performance heavily depends on the accuracy and completeness of the input data [39]. In real-world applications, errors in entering ABG values or missing critical context, such as comorbidities, medications, or clinical history, can lead to inappropriate or incomplete recommendations. Furthermore, ChatGPT cannot independently verify or question the validity of the provided data and lacks the ability to integrate this information into a broader clinical framework [40]. For instance, ABG interpretation often relies on understanding the patient’s clinical trajectory, underlying conditions, and current treatment plans [5]. These nuances, which human clinicians use to refine their decisions, are beyond ChatGPT’s capabilities. Over-reliance on AI systems without adequate human supervision can lead to errors and serious consequences [41]. Additionally, concerns about data privacy and security in AI-based systems must be addressed before widespread implementation [42].
While this study offers encouraging results regarding the use of ChatGPT in ABG interpretation, further research is warranted to explore its applicability across different clinical scenarios and healthcare settings. Future investigations may also provide a deeper understanding of its role in clinical decision-making and integration into healthcare workflow processes.
This study demonstrates that ChatGPT performs comparably to clinicians in ABG interpretation and treatment planning, offering significant potential as a decision-support tool in anesthesia and critical care. By enhancing clinical accuracy and reducing cognitive workload, ChatGPT could play an increasingly valuable role in modern healthcare. However, its use should remain complementary to clinical expertise, emphasizing the importance of validation and ethical considerations in its application.
The data presented in this paper are available upon reasonable request from the corresponding author.
USÇ—development of methodology, conception, and study design; data collection, analysis, interpretation, drafting of the manuscript, critical revision, and final approval. BA—development of methodology, conception and study design, analysis, statistics critical revision and final approval. HKÖ—development of methodology, analysis, and final approval. HY—development of methodology, analysis, and final approval. ÜK—conception and study design, critical revision, and final approval. All the authors have read and approved the final version of this manuscript.
This study complied with the principles of the 1964 Declaration of Helsinki and its amendments. Patient data were anonymized to ensure confidentiality. Ethical approval was obtained from the Institutional Review Board on 02 January 2025 (Approval Number: 02.01.2025-9-285) SBU Adana City Education and Research Hospital. Additionally, all participating physicians provided written informed consent prior to their inclusion in the study.
Not applicable.
This research received no external funding.
The authors declare no conflict of interest.