從基因名稱獲取蛋白序列並完成 BLAST 比對
案例演示 查詢經過審閱的人血紅蛋白 α 亞基條目並獲取序列
從人 HBA1 基因名稱出發,找到已審閱的 UniProt 蛋白及標準 FASTA 序列,提交 BLAST 檢索並檢查完成後的報告。本例使用已知蛋白,演示序列獲取與比對方法。
開始前,按科學資料庫啟用所需 Connector,選擇已連線的模型,並確保 Notebook 執行環境可用。
1. 找到蛋白並獲取 FASTA 序列
還不知道登入號時,Genes & Ontologies 可先發現 UniProt 條目。向 search_uniprot_entries 提供基因名、蛋白名稱短語或物種。organism_id 匹配指定分類單元;reviewed: true 選擇 Swiss-Prot,false 選擇未審閱的 TrEMBL 條目,省略則包含兩者。續查時使用 next_cursor,保持篩選條件和每頁大小不變。
啟用 Genes & Ontologies,開啟已連線模型且 Notebook 執行環境可用的會話,傳送:
Use Genes & Ontologies through Session Notebook. Call search_uniprot_entries
with gene HBA1, organism_id 9606, reviewed true and page_size 5.
Save the query and complete response as hba1-uniprot.json. Check the
returned organism and protein identity, then retrieve the canonical FASTA
for the matching accession with get_uniprot_entries. Add that response to
the JSON and save hba1-query.fasta. Keep everything in English. Preserve
missing or empty results; do not invent sequences.
使用 FASTA 前先開啟 JSON。本次返回 P69905 / HBA_HUMAN,物種為 Homo sapiens,長度 142 個氨基酸,基因名稱包含 HBA1 和 HBA2。響應標明 UniProt 版本 2026_03、total_results: 1 和 has_more: false。FASTA 標題行保留登入號和物種,序列包含 142 個殘基。按基因名查詢可能返回關聯多個基因的蛋白條目,不能據此假定基因與條目一一對應。

UniProt 查詢與 FASTA 響應 · 標準 FASTA 序列
2. 提交併跟蹤 BLAST 任務
需要檢索相似序列時,啟用 Genomes,使用其中的三個 BLAST 操作。序列會傳送至公共 NCBI 服務,請使用公開或已獲授權的輸入。
- 使用序列、
molecule_type和相容資料庫呼叫一次blast_submit,儲存返回的rid與輪詢提示。對於上面的蛋白,molecule_type: protein和database: swissprot表示進行蛋白檢索。 - 針對該 RID 呼叫
blast_status。同一 RID 的請求至少間隔 60 秒,所有 BLAST 請求至少間隔 10 秒;服務要求更長等待時,按更長間隔執行。WAITING表示仍在排隊或執行,應保留 RID 繼續查詢,不要重複提交。 - 出現
READY後,仍按同樣的間隔呼叫blast_results。可選格式為json2、xml2、text和tabular。報告大小上限為 2 MiB,必要時減少命中數量。表格格式可能帶註釋,不能直接當作 CSV 表。 - 從實際報告核對查詢長度、實際資料庫、命中登入號、比對範圍、相同殘基比例和 E-value。序列相似性本身不能證明功能;已知血紅蛋白序列適合學習操作,不能當作發現未知蛋白的案例。
如果提交返回 blast_submission_unknown,表示無法確定是否已被接受,不要自動重複提交,應保留響應和已有 RID。提交回執或 WAITING 狀態都不是完成後的比對結果。具體輸入和返回條件見 BLAST 參考。
3. 開啟並解讀完成後的報告
在上面的同一會話中繼續蛋白序列案例。保留提交回執,後續請求才能接著查詢同一個任務。傳送:
Continue with the public P69905 FASTA retrieved above. Use Genomes through
Session Notebook. If this session already has a BLAST RID, resume that RID;
otherwise call blast_submit once with molecule_type protein, database
swissprot and hitlist_size 5, then save the receipt. Space requests for the
same RID by at least 60 seconds and follow any longer server delay. Check
blast_status; if it is still WAITING, keep the RID for a later check rather
than submitting again. After READY, wait the required interval and retrieve
blast_results in json2 format. Save hba1-blast-raw.json, hba1-blast-hits.csv
and hba1-blast-results.md. Include the actual database, query length,
accessions, alignment coordinates, identity counts and E-values. Explain
the coverage and identity calculations. Keep everything in English and
preserve actual errors or empty results. Never invent alignments.

報告就緒後,開啟 hba1-blast-results.md,把結果表與 hba1-blast-raw.json 對照。本例報告記錄的程式為 BLASTP 2.17.0+,實際資料庫為 swissprot,查詢序列長 142 個氨基酸,返回 5 個命中:
| 登入號 | 相同殘基數 / 比對長度 | 查詢覆蓋率 | E-value |
|---|---|---|---|
| P69905 | 142/142 (100%) | 100% | 1.99033e-100 |
| P01923 | 140/141 (99.29%) | 99.30% | 1.06845e-98 |
| Q9TS35 | 140/142 (98.59%) | 100% | 2.38346e-98 |
| P06635 | 139/142 (97.89%) | 100% | 3.57742e-98 |
| P01924 | 138/141 (97.87%) | 99.30% | 3.00466e-97 |

這裡按每個命中的第一個 HSP 計算:一致率是相同殘基數除以比對長度;查詢覆蓋率是包含首尾位置的查詢跨度除以 142。以 P01923 為例,查詢位置為 2–142,因此覆蓋率為 141/142 = 99.30%,一致率則為 140/141 = 99.29%。兩者回答的問題不同,也都不是功能判斷正確的機率。
完整結果報告 · 五個命中的結果表 · NCBI JSON2 報告
第一項 P69905 就是輸入序列本身,100% 的一致率和覆蓋率用於核對已知序列。其他命中展示序列相似性,不代表發現了新功能。保留原始報告和查詢序列;資料庫更新後,命中列表可能變化。