GRADE 方法學概述與維生素 D 互動實務案例(Overview of GRADE Methodology and Interactive Vitamin D Practice Example)
- 會議/場次: ENDO 2026/Endocrine Society Guideline Methodology: A Primer for Users and Aspiring Panelists(PD12)
- 短講: PD12-GRADE
- 講者: Spyridoula Maraka, MD, MS
整理稿
從臨床問題、證據走到推薦
Spyridoula Maraka 接續 Christopher McCartney 的歷史回顧,說明臨床實務指南可簡化為一條路徑:先提出臨床問題,再評估最佳可得證據,最後形成包含推薦的陳述。GRADE(Grading of Recommendations Assessment, Development and Evaluation)是評估證據確定性並制定醫療推薦的結構化方法;其中 Evidence-to-Decision(EtD)framework 讓 panel 在做決定時,不只看效益與傷害,也系統性考量價值觀、資源、公平性、可接受性與可行性。
Endocrine Society 認為,忠實遵循 GRADE 可減少 guideline-development process 中的偏差與扭曲、提高透明度,並支持符合 Institute of Medicine/National Academy of Medicine 的標準。GRADE 的重點不是把不確定性藏起來,而是清楚呈現證據如何被評估、各項判斷如何連到推薦。
從主題式長篇敘述轉向聚焦問題
舊式指南常同時回顧生理、病理、鑑別診斷、相關藥物與一般臨床評估;現行流程則把問題、systematic review、EtD 與推薦直接串接。因此,指南通常會有較少的推薦、較有限的寫作任務與較短的篇幅,但臨床使用時更清楚、也更容易追蹤推薦依據。
投影片以非糖尿病患者低血糖指南為例:與其概述整個主題,新格式會聚焦於「嚴重低血糖門診患者應使用需現場調配的 glucagon,或不需調配的 glucagon 製劑」這類可採取行動的決策問題,再連到 evidence appraisal、EtD、推薦理由與 research considerations。
Guideline Development Panel 與 COI 管理
Guideline Development Panel(GDP)依主題納入成人與兒童內分泌、一般內科、婦產科、營養或流行病學等 subject-matter experts;另有兩位 methodologists,一位主導 GRADE implementation,另一位主導 systematic review 與 meta-analysis。現行 panel 至少有一位 patient representative;若指南預計由其他學會 endorsement,也會納入相關學會代表。
Conflict-of-interest(COI)政策貫穿 planning and scoping、panel selection、evidence review、recommendation development、public review and comment、publication and updates 六個階段。講者舉例,guideline chairs 必須完全沒有 COI;其他成員的揭露與管理則一路持續到發表與更新。這些程序是為了讓決策以證據為基礎,而非受個人利益左右。
Background、foreground 與 PICO 問題
Background question 是對健康狀況的廣泛問題,例如 MRI 診斷膝骨關節炎的準確度;它適合提供知識與引導後續思考,但不一定能直接回答眼前的臨床決策。Foreground question 則鎖定特定族群與決策,例如某族群是否應使用 drug X 而非 drug Y,能直接導向可操作的推薦。
Foreground question 以 PICO 定義:P 是 population,需界定年齡、性別、共病與照護情境;I 是 intervention,需說清楚藥物、計畫或檢查;C 是 comparator,可能是無 active treatment 或另一種治療;O 是能影響決策的 patient-important outcomes,例如死亡、疾病負擔與生活品質。Panel 會先廣泛腦力激盪,再按資源預先限定問題數量,通常收斂到約 10 個。GRADEpro/GDT 畫面顯示成員以 1–9 分排序優先度,再決定問題是進入下一階段、僅列於 introduction,或排除。
選擇真正重要的 outcomes
每個問題通常選 3–5 個 outcomes,最多不超過 7 個,並同時納入 desirable 與 undesirable outcomes。選擇應由「對病人和決策是否重要」驅動,不能因預期找不到研究就先排除;缺乏證據本身可能揭示 research gap。若 critical outcome 沒有直接資料,panel 可以考慮 surrogate outcome,例如以 bone mineral density 代替 fracture,但這會因 indirectness 降低證據確定性。
Outcomes 同樣以 1–9 分分級:7–9 分為 critical for decision-making,4–6 分為 important but not critical,1–3 分為 low importance。講者用 mortality、fractures、cardiovascular disease、cancer 與 adverse events 示範排序,也提醒:「不是所有重要的事都被測量,而被測量的事也不一定都對決策重要。」
Systematic review 與證據確定性
Evidence-synthesis team 依 PICO 進行 systematic review 與 meta-analysis,將研究整合為 effect estimate,並為每一 outcome 建立 evidence profile。證據確定性由五個領域判斷:risk of bias、inconsistency、indirectness、imprecision、publication bias。
GRADE 將 certainty of evidence 分為 high、moderate、low、very low。High 表示很有信心 true effect 接近估計值;moderate 表示可能接近,但仍可能有實質差異;low 表示信心有限;very low 則表示真實效果很可能與估計值有重大差異。Maraka 指出,內分泌領域常得到 low 或 very low certainty,panel 必須在這種不完美的證據下做出透明判斷。
EtD 與推薦方向、強度
GDP 接著使用 GRADEpro 的 EtD 表格,綜合 desirable effects、undesirable effects、certainty、values、resources、cost effectiveness、equity、acceptability 與 feasibility,決定推薦的方向與強度。若效益與傷害的權衡清楚,可形成 strong recommendation,使用「we recommend」;多數人會選擇該處置,臨床人員在幾乎所有相應情境下都應提供,政策也通常可採納。
若權衡不清楚,則形成 conditional recommendation,使用「we suggest」。多數人可能選擇該處置,但仍有相當比例不會,因此醫療人員必須考慮個別因素、協助病人依自身價值觀做 shared decision-making;政策面也需要更多 stakeholder 討論。Panel 的選項包括 strong/conditional recommendation for or against intervention,也可能在兩者權衡相當或不確定時,給出 intervention 或 comparison 皆可的 conditional recommendation。
75 歲以上成人的 vitamin D 互動案例
Maraka 邀請現場聽眾暫時扮演沒有 COI 的 guideline panel,並以 Endocrine Society 新近發表的 vitamin D 指南示範。PICO 問題是:75 歲以上成人是否應在 recommended dietary intake 之外,接受 empiric vitamin D supplementation,而非不做 empiric supplementation?
聽眾以文字雲提出 mortality、fractures、falls、quality of life、frailty、cardiometabolic disease 與 dementia 等 outcomes。正式 GDP 將 fractures、falls、respiratory tract infections、all-cause mortality 列為 critical;將 nephrolithiasis 與 kidney disease/renal failure 兩項 adverse events 列為 important。
Vitamin D 證據表如何閱讀
Summary of Findings 顯示:任何 fracture 為 15 項 RCT、43,585 人,RR 1.01(95% CI 0.94–1.08),high certainty;falls 為 16 項 RCT、12,342 人,RR 0.97(0.91–1.03),moderate certainty;respiratory tract infections 為 1 項 RCT、821 人,HR 1.11(0.94–1.30),low certainty。
All-cause mortality 為 25 項 RCT、49,879 人,RR 0.96(0.93–1.00),high certainty;點估計顯示很小的風險下降,相當於每千人由 158 降至 152,而信賴區間上限到達 null effect。Nephrolithiasis 為 3 項 RCT、6,306 人,RR 0.94(0.54–1.65);kidney disease/renal failure 為 3 項 RCT、5,634 人,RR 0.76(0.44–1.32),兩者皆為 moderate certainty,且區間寬,不能把點估計解讀為確定的保護效果。Falls 因 inconsistency 降級;respiratory infection 的評估同時有 risk-of-bias concerns 與 imprecision。
Slides 延伸補充:效果方向必須連同 95% CI、absolute effect 與 certainty 一起閱讀;「點估計稍高或稍低」不等於已證明增加或降低風險。
講者最後說明,對一個 guideline question 做整體 certainty 判斷時,通常取 critical outcomes 中最低的 certainty。因此,即使 fractures 與 mortality 是 high certainty,只要 critical 的 respiratory infection 是 low certainty,整體問題通常仍會落在 low certainty。這段短講到 evidence synthesis 為止;接著由 Juan P. Brito 接手 EtD 的進一步討論,屬於下一個正式 presentation。