Identifying Diabetes-Relevant Gene Programs across Metabolic Tissues in the Common Metabolic Diseases Knowledge Portal
ADA2026

Identifying Diabetes-Relevant Gene Programs across Metabolic Tissues in the Common Metabolic Diseases Knowledge Portal

2026-06-05

Identifying Diabetes-Relevant Gene Programs across Metabolic Tissues in the Common Metabolic Diseases Knowledge Portal

開場:從「我關心的基因」走向細胞狀態、基因程式與疾病關聯

Julie Jurgens, PhD 代表 Knowledge Portal Network 介紹 Common Metabolic Diseases Knowledge Portal(CMDKP, cmdkp.org)中的新功能。她在 Broad Institute 與 Noël Burtt 團隊合作,所屬為 Burtt Lab / Knowledge Portal Network;本場演講無利益衝突聲明,現場允許拍照。講題正式名稱為「Identifying Diabetes-Relevant Gene Programs across Metabolic Tissues in the Common Metabolic Diseases Knowledge Portal」,機構標示為 Broad Institute of MIT and Harvard。

這場演講的核心問題是:如果研究者曾經想問——「我最關心的基因參與哪些細胞狀態(cell states)與基因程式(gene programs)?這些又如何與疾病相關?」——CMDKP 現在提供了一個新的開放平台,讓使用者可以直接探索這些問題。這個新平台不是單純顯示基因在哪裡表現,而是把單細胞資料、已知細胞狀態、資料驅動推論出的基因程式,以及人類遺傳學支持的疾病性狀連結在一起。

CMDKP 的定位:把遺傳訊號連到常見代謝疾病

Common Metabolic Diseases Knowledge Portal(CMDKP) 是一個開放、社群驅動的網頁入口網站,目標是把人類遺傳訊號連結到各類常見代謝疾病。過去 CMDKP 的核心工作主要建立在 genetic and genomic data 上:彙整多個遺傳資料集、進行跨資料集 meta-analysis,並搭配 functional annotations、社群整理的專家知識,以及多種 bioinformatic analyses,讓研究者能問出幾類常見問題:

  • 我研究的 phenotype,其風險牽涉哪些基因?
  • 我正在研究的 gene 是否在人類疾病中扮演角色?
  • 某個 region、variant、phenotype 或 tissue 的資料,是否能支持特定疾病機轉?

CMDKP 的首頁定位是「providing data and tools to promote understanding and treatment of common metabolic diseases」,可用 gene、region、variant、phenotype 或 tissue 搜尋;範例包括 PCSK9 gene、PCSK9 region、rs1260326,或 chr9:21,940,000–22,190,000 這類基因座座標。

CMDKP 涵蓋的疾病範圍並不只限於糖尿病本身,而是包括多種共同代謝疾病。平台疾病範疇包含 T1D/T2D 與 prediabetes、obesity、kidney disease、liver disease、heart disease;整體資料規模標示為超過 1,000 個 phenotypes、超過 900 個 genetic datasets、超過 7,000 個 disease-relevant tissues 中的 genomic annotations,並整合約 10 種 bioinformatic methods,其中包含新方法與社群專家知識。

從遺傳與基因體資料延伸到單細胞解析度

CMDKP 近期的擴充重點,是讓使用者能以細胞解析度(cellular resolution)探索自己關心的疾病與基因。這是透過 Single Cell Browser 完成的:平台整合了來自多個來源、涵蓋 9 個代謝相關人類組織的單細胞圖譜,有些資料集規模超過百萬細胞。這些單細胞圖譜由 Kyle Gaulton 團隊建立,並由 Knowledge Portal 團隊整合進 CMDKP。

使用者可以到 CMDKP 的 KP Labs > Single Cell Browser 進入功能頁面。Single Cell Browser 提供互動式探索介面,包括:

  • cell type clustering 的 UMAP plots;
  • 查詢感興趣基因,觀察其在哪些細胞類型中表現;
  • 各 cell type 在資料集中的比例;
  • gene expression 的 violin plots;
  • 每個組織中各 cell type 的 marker gene expression。

當時 Single Cell Browser 整合的 9 個代謝相關人類組織及資料規模為:artery,7 個 sources、43 位 donors、241,983 cells;heart,8 sources、108 donors、1,384,590 cells;hypothalamus,2 sources、14 donors、567,691 cells;kidney,7 sources、155 donors、316,161 cells;liver,10 sources、178 donors、1,058,446 cells;skeletal muscle,7 sources、183 donors、364,926 cells;pancreas,4 sources、327 donors、1,549,550 cells;white adipose subcutaneous,5 sources、752 donors、1,685,704 cells;white adipose visceral,4 sources、38 donors、202,415 cells。

這個瀏覽器已經能幫助研究者回答「基因在哪些細胞類型表現」這類問題;接下來的新功能,則進一步把問題推向:基因是否屬於某些協同表現的基因程式?這些程式是否對應到已知細胞狀態?它們是否與人類代謝疾病性狀有遺傳學支持?

新功能:以 NMF 進行 gene program inference

新的分析功能是 gene program inference。這裡使用的是 non-negative matrix factorization(NMF) 這類技術。NMF 是一種 dimensionality reduction technique,用來在高維資料中找出潛在結構(latent structures),也就是所謂的 factors。概念上,原本很大的矩陣可以被分解成較小的子矩陣,研究者便可透過這些 factors 找到原始資料中不容易直接看出的隱藏模式。

NMF 的數學直覺是將 original matrix 近似分解為 feature matrix 與 coefficient matrix;這類方法可用於 recommender system、biological feature analysis、image processing 與 clustering problem。

Julie 用 Netflix 的電影推薦作為類比。假設一個人像她一樣對 Netflix 上癮,每次想取消訂閱,Netflix 就推送「Love Is Blind 3 現在可以看了」這類精準推薦,讓人又繼續訂閱。Netflix 如何知道使用者可能想看什麼?一種方式就是從「使用者 × 影片」的矩陣開始:矩陣中記錄哪些使用者看了哪些電影或節目。接著透過矩陣分解,把這個 user-item interaction matrix 拆成 user matrix 與 item/movie matrix,並由 NMF 推導出 factors;在這個例子中,factors 對應的就是使用者偏好,例如偏好特定類型的節目或電影。

CMDKP 將同樣的想法應用於單細胞基因表現資料。不再是 users 與 movies,而是 cells × genes 的矩陣。平台使用 LIGER 方法推導 factors;在這裡,factors 就是協同表現的基因模式,也就是 gene programs

Gene programs 是一群具有協調表現模式的基因,通常反映某種底層生物過程、細胞狀態或調控機制。換句話說,平台先從 cell-by-gene matrix 出發,再把資料分解成 gene matrix 與 cell matrix,進而 computationally infer gene programs,協助使用者探索可能的重要生物過程。

在 CMDKP 的 NMF 概念圖中,輸入資料被描述為 log of read count matrix,且 NMF 要求矩陣值為 non-negative;分解後得到的 gene matrix 與 cell matrix 共同定義出協同表現的 gene programs。

Gene programs 與 cell states:資料驅動推論加上已知生物學

完成 gene program 分析後,團隊意識到這類方法的優點是非常 exploratory:它能讓使用者從資料中找出事先可能不知道的模式。然而,這同時也可能成為限制,因為資料驅動的因素不一定容易解釋,也不一定能直接對應到已知生物學。

因此,CMDKP 把 gene program analyses 搭配另一種互補方法:cell states

這裡的 cell states 指的是 transcriptionally defined cell populations,也就是根據轉錄體特徵定義出的細胞狀態,並且從已發表文獻與其他資料來源中整理、curate 出來,用來捕捉已知生物學。這些不是完全由演算法「盲目」推論的 factors,而是更接近由 marker genes 與文獻知識支撐的 biological states。

因此,平台同時提供兩種視角:

  • Gene programs:由資料驅動、computationally inferred 的協同基因表現模式。
  • Cell states:由文獻與已知 marker-defined biology curate 出來的細胞狀態。

兩者結合後,使用者可以更完整地理解某個 cell type 如何與人類疾病相關。平台將 gene programs 定義為「groups of genes with coordinated expression patterns in a cell type」,而 cell states 則定義為「transcriptionally defined cell populations curated from published literature and capturing known biology」;cell state 的例子可包括由 ER stress marker genes 定義出的功能狀態。

CMDKP 與既有資源的差異

現有許多資源通常提供兩類功能:

  • single-cell atlas exploration;
  • post-GWAS cell-type 或 gene-set enrichment analyses。

CMDKP 的新功能與這些資源不同之處,在於它不只是提供 atlas 瀏覽或 GWAS 後 enrichment,而是把多層資料系統性地連接起來:

  1. harmonized healthy and disease single-cell maps;
  2. key common metabolic tissues;
  3. curated cell states;
  4. informatically inferred gene programs across multiple tissues at scale;
  5. cell states 與 gene programs 的 systematic comparison;
  6. 回到人類遺傳學根基,提供 metabolic traits and diseases 的 genetic support。

所有分析都可以透過 CMDKP 公開瀏覽。這個功能當時的整體規模標示為 9 個 tissues、105 個 cell-type groups、598 個 curated states、1,398 個 inferred factors;分析以 precomputed analyses 與 visualizations 形式公開提供,使用者不需要自行重新跑完整流程即可探索。

Cell State & Program Explorer:以 PNPLA3、liver hepatocyte 為例

新功能可從 CMDKP 的 KP Labs > Cell State & Program Explorer 進入。進入介面後,使用者先輸入感興趣的 gene,接著選擇 tissue 與 cell type。Julie 以 PNPLA3 作為例子,選擇 liverhepatocyte

Cell State & Program Explorer 的功能描述是:比較基因在 cell types、curated cell states 與 computationally inferred gene programs 中的 expression,並提供 genetically supported links to human traits,以揭示已知與潛在的新生物學。

在 PNPLA3 的範例中,介面會先回答「PNPLA3 在哪裡表現?」並列出不同 tissues 與 liver 內的 cell types。在 liver 的 cell type 清單中,hepatocyte 是主要相關細胞類型;畫面示例中 hepatocyte 的 expression 指標為 ABS -0.17、SPEC 5.12,而其他如 Schwann cell、cholangiocyte、hepatic stellate cell、endothelial、macrophage、mast cell、NK cell、B cell、T cell 的 specific expression 則較低,顯示 PNPLA3 在 liver hepatocyte context 中較具特異性。

左側:curated cell states

進入 liver hepatocyte 後,介面左側顯示 cell states。這些 cell states 是已 curated、marker-defined 的 biological pieces,與 liver hepatocyte 中已知生物狀態或功能程式相關。表格會呈現每個 state 的 expression,並用 absolute(ABS)specific(SPEC) 兩種值表達。

使用者可將滑鼠移到欄位或列上,查看這些定義的含義、該 cell state 的 description,以及 marker genes。點擊任一列後,會進入更詳細的 metadata,包括:

  • 這個 state 代表什麼;
  • 它如何被 curated;
  • 其 provenance;
  • 使用的 references;
  • 與哪些 related programs 有關。

PNPLA3 / liver / hepatocyte 範例中列出的 6 個 curated cell states 包括 Periportal Hepatocyte、Bile Acid Transport Metabolism、Pericentral Hepatocyte、Proliferating Hepatocyte、Hepatocyte Identity,以及 Acute Phase Inflammatory Response 相關狀態;範例中的 expression 指標大致落在 ABS 約 -0.15 至 -0.21、SPEC 約 0.12 至 -0.30。

Bile Acid Transport Metabolism 為例,它被定義為 hepatocyte 中的 functional activity program,反映 bile acid transport metabolism。其 marker genes 包括 ABCB11、ABCC2、BAAT、CYP7A1、CYP8B1、HSD3B7、NR1H4、SLC10A1、SLCO1B1、SLCO1B3;state ID 為 liver_hepatocyte_bile_acid_transport_metabolism,curation version 標示為 2026.06.04,參考文獻包含 Aizarani et al., Nature 2019。

這類功能性 cell state 需要謹慎解讀。介面中的詮釋提醒指出:functional states often gradients,因此若某個 gene 在此 state enrichment,應解讀為與較高功能活動相關,而不是代表該 gene 屬於互斥的離散細胞亞型;下一步應檢查該 gene 是否也在相關 functional programs 或整體 cell-type expression 中 enriched,不應把它解讀為 mutually exclusive cell subtype。

右側:inferred gene programs

介面右側則是 gene programs。這些 gene programs 是 data-driven、computationally inferred 的 latent factors。操作方式與 cell states 類似:使用者可以 hover 查看基本資訊,也可以點擊某個 row 取得更完整內容。

點擊 gene program 後,頁面會提供:

  • 該 program 的 detailed explanation;
  • program 中的 top genes;
  • 它與哪些 curated states 相符;
  • 它與哪些 human traits 或 disease traits 有關聯或 enrichment。

這裡的重要設計是:gene programs 並不是孤立呈現,而是被放在 curated cell states 旁邊,讓使用者能比較「資料推論出的模式」與「文獻/marker 支撐的已知狀態」是否一致,或是否揭示出新的生物學。

Cell states 與 gene programs 的關係:Spearman correlation 與探索性檢查

在同一個介面的下方,可以探索 cell states 與 gene programs 之間的關係。平台以 Spearman correlation 建立兩者之間的相關性 heatmap;在 Julie 展示的視圖中,較低相關性以紅色表示,較高相關性以藍色表示。

這部分目前仍偏 exploratory。Julie 強調,就她所知,這類在大規模多組織資料中系統性比較 curated cell states 與 inferred gene programs 的分析尚未廣泛完成,因此團隊希望研究社群實際使用、測試,看看是否能找到預期中的連結,也可能發現未預期的新關係。

團隊自己已進行一些 spot checks,其中令人放心的一個例子是:hepatocyte identity 的 cell state 與對應的 hepatocyte identity gene program 呈正相關。這表示至少在某些已知生物學場景中,資料驅動推論出的 gene program 能與 curated known biology 對應。

使用者也可以切換不同 correlation 視圖。例如,若想檢查 cell state 中 marker genes 與 gene programs 的關係,可以從下拉選單選取相應選項,介面會顯示對應 heatmap。這使得分析不只停留在 state-program 整體層級,也能往 marker gene set 與 program gene set 之間的關係延伸。

連回人類遺傳學:以 PIGEON 建立 traits 與疾病支持

介面最下方提供 cell states 與 gene programs 到 human traits 的 genetic linkages。這部分使用的方法稱為 PIGEON,全名為 Priors Inferred from Gene Annotations。這是一種 statistical genetic method,由 Jason Flannick 團隊發展。

PIGEON 的概念是:輸入一組 genes,例如某個 gene program 中的 genes,或某個 cell state 中的 marker genes,然後把這組 genes 與人類性狀或疾病的遺傳資料連結起來,評估這些 gene sets 是否與特定 human traits 有遺傳支持。

在 PNPLA3 / liver hepatocyte 的例子中,某些 programs 與 lipid traits 呈現高度相關,這與既有文獻中 PNPLA3 和肝臟脂質代謝相關的觀察相符。也就是說,平台不只是顯示 PNPLA3 在 hepatocyte 表現,也能把 hepatocyte 中的 gene programs、curated functional states 與 lipid traits 的遺傳學證據放在一起,幫助研究者形成更具機轉性的假說。

在 metadata 設計中,human genetic trait anchors 可包含 joint beta、marginal beta 與 method 等欄位;cell state 的詳細頁也可列出 related programs with GSEA P < 0.05,例如 Bile Acid Transport Metabolism state 可連到 Xenobiotic Lipid Metabolism 等相關 program。

總結:CMDKP 的新資源如何支持糖尿病與代謝疾病研究

Gene program 與 cell state 分析是理解 cellular function 及其 disease roles 的互補方法。Gene programs 提供資料驅動、探索性的協同基因表現模式;cell states 則提供基於文獻與 marker-defined biology 的已知狀態。CMDKP 將這兩者放在同一個平台中,並與 human genetics 連結,使研究者能在單一介面中探索:

  • harmonized healthy and disease single-cell maps;
  • key common metabolic tissues;
  • curated cell states;
  • informatically inferred gene programs across tissues at scale;
  • cell states 與 gene programs 的 systematic comparisons;
  • metabolic traits and diseases 的 genetic support。

所有分析皆透過 CMDKP 公開提供,並以可瀏覽、可互動的方式呈現。Julie 將此稱為 first-of-a-kind resource,目標是讓研究者不必自行整合大型單細胞與遺傳資料,就能直接探索疾病相關基因程式與細胞狀態。

NIDDK Data Science Corner 與相關開放資源

這項新功能只是 NIDDK 資助的多個 open-access resources 之一。在 ADA 現場的 exhibit hall 有 NIDDK Data Science Corner,Julie 邀請聽眾前往了解相關資源。她特別提到團隊也與其他 groups 共同開發:

  • Mammalian Adipose Tissue Knowledge Portal(MATKP)
  • Pankbase,也就是前一位講者 Fan Feng 介紹過的 pancreas knowledge base

NIDDK Data Science Corner 中同時列出的資源還包括 Common Metabolic Diseases Knowledge Portal(cmdkp.org)、Mammalian Adipose Tissue Knowledge Portal(matkp.org)、Pankbase(pankbase.org)、dkNET(dknet.org)、Integrated Islet Distribution Program(iidp.coh.org)、Mouse Metabolic Phenotyping Centers(mmpc.org)、Type 1 Diabetes TrialNet(trialnet.org)、Rare and Atypical Diabetes Network / RADIANT(atypicaldiabetesnetwork.org)。

致謝

Julie 感謝整個團隊讓此資源得以完成,也感謝 NIDDK 與 FNIH 的 funding,以及與會者參與。

團隊成員與角色包括:DK Jang(UX/UI、principal design architect)、MacKenzie Brandes(project management)、Lucas Nguyen(computational biology)、Trang Nguyen(computational biology)、Alex Shilin(UX/UI、design engineering)、Quy Hoang(software engineering)、Patrick Smadbeck(software engineering)、Drew Hite(software engineering)、Parul Kutarkar(computational biology)、Luca Tucciarone(computational biology)、Yuna Lee(computational biology)、Weston Elison(computational biology),以及 Ben Voight、Kyle Gaulton、Noël Burtt、Jason Flannick。Funding 來源包括 National Institute of Diabetes and Digestive and Kidney Diseases(NIDDK)UM1DK105554 與 Foundation for the National Institutes of Health。

討論與問答

Pancreas / pancreatic biology 是否也已檢查?

第一個問題指出,Julie 展示的例子主要與 liver 或 hepatic regulation 相關,詢問是否已經檢查過 pancreas 或 pancreatic biology 相關例子。

Julie 回答,平台目前已經有 9 個組織的 maps,因此 pancreas 也在範圍內。團隊已經對一些 genes 做了 spot checks,但更系統性的分析仍在進行中。她也邀請社群提出特別值得測試的 genes,尤其是那些「生物學已知,但還不是非常廣為人知」的案例。這類案例很適合檢查 gene programs 是否能正確推論出相關 biology,並可能進一步提供 mechanistic insight。

社群是否能像 Wikipedia 一樣協助改善資料品質?

第二個問題來自一位已經在聽講時試用資源的聽眾。問題是:平台是否有機制讓社群貢獻資料品質,例如提交人工 curated 的 Wikipedia-style updates?提問者也指出,這會需要平台端有額外 bandwidth。

Julie 認為這是很好的想法。CMDKP 的價值很大程度來自社群參與與專家知識,因此她希望能更有效地利用研究社群的 expertise。現階段,如果使用者有建議,可以寄信到 help@cmdkp.org,團隊會透過 help desk 與使用者連結。她也表示,團隊應該考慮建立更系統性的表單或提交機制,讓社群能更容易提供更新與改進建議。

Take-home Summary

  • CMDKP 是開放、社群驅動的平台,用來把 genetic signals 連結到 common metabolic diseases。
  • 新增的 Single Cell BrowserCell State & Program Explorer 讓研究者能在單細胞解析度下探索基因、cell types、cell states、gene programs 與 traits。
  • Gene programs 由 NMF/LIGER 等資料驅動方法推論,代表協同基因表現模式;cell states 則由文獻與 marker-defined biology curated 而來。
  • 將 gene programs 與 cell states 並列比較,可檢查資料推論出的 latent factors 是否對應到已知生物學,也可找出新的疾病相關機轉。
  • 平台進一步用 PIGEON 將 gene sets 連回 human traits,提供 metabolic traits and diseases 的 genetic support。
  • 這個資源已公開於 cmdkp.org,並鼓勵研究社群測試、回饋與共同改善。