新着研究 / Computer Vision / Generative AI
病変を色で示す画像編集型の領域分割、未知の撮影条件で精度を検証
InstEditSeg
InstEditSegの原論文から掲載した図です。細部はクリックして拡大できます。
原論文の図の説明を読む
Figure 1: Conceptual illustration of the proposed framework. Traditional methods generate binary masks. Our method reformulates segmentation as an image editing task. It renders color-coded regions on the original image to align with generative priors and mitigate domain shift.
概要
- 元画像に色を重ねる画像編集として病変の領域分割を行い、指示文で対象を指定する。
- 未学習のPolypGenではDice83.92%、未学習のISIC2017ではDice83.14%を報告した。
- 色指定への感度と、属性を指定した病変の選択が未対応である点が制約となる。
補足:編集した模式図で仕組みを確認する
VISUAL EXPLAINER
図でつかむ、InstEditSegの仕組み
- 01医療画像と指示文
- 02画像特徴を抽出
- 03指示に沿って着色
- 04病変を重ねた画像
原論文の説明をもとにした模式図です。処理の細部は省略しています。
研究の背景
ポリープや皮膚病変の輪郭を画像上で特定する作業は、観察や診断を支える。ところが、病変と周囲の色の差が小さく、境界がぼけ、照明や撮影機器も変わるため、自動の領域分割は難しい。
従来法の多くは病変だけを示す白黒の画像を出す。著者らは、元画像を残した着色画像を作る方法が、学習済みの画像生成モデルを医療画像へ使ううえで有効かを調べた。
手法
学習時は、正解の病変領域にランダムな色を割り当て、元画像へ重ねた編集後画像を用意する。「ポリープを赤で示す」といった指示文も組み合わせる。拡散モデルとは、画像へ加えたノイズを段階的に取り除くことで画像を作る方式で、この方法では元画像と指示文を条件に編集後画像を復元する。
境界を捉える補助として、画像の特徴を抽出するDINOv3を使う。その特徴を粗い構造から細部まで複数の大きさに変換し、生成を担うU-Netへ渡す。推論時は元画像と指示文から着色画像を生成し、各ノイズ除去段階では指示文あり・なしの2通りを計算して文の効き方を調整する。
新規性
著者らの提案は、ポリープと皮膚病変の領域分割を、元画像に指定色を重ねる画像編集として統一した点にある。学習時には正解領域の色と指示文を変え、推論時には画像と文だけから着色画像を作る。
さらにDINOv3の多段階の画像特徴を生成過程へ注入し、境界の手掛かりを補う。病変の種類ごとに専用の出力部を設けない構成も特徴だ。
従来手法との違い
比較対象のU-NetやEMCADなどは、各画素が病変かどうかを示す二値マスクを直接予測する。InstEditSegは病変を指定色で塗った画像を生成するため、出力の形式と学習目標が異なる。
著者らの表では、既知の検査画像でEMCADなどが上回る条件もある。一方、学習に使わなかったPolypGenではInstEditSegのDiceとIoUが比較法中で最高だった。優位性は全データに一様ではない。
実験結果
著者らはポリープ4種の学習用データをまとめ、各試験分と未学習のPolypGenで評価した。皮膚病変はISIC2016で学習し、ISIC2017を未学習の試験データとした。Diceは予測領域と正解領域の重なり具合、IoUは両者の共通部分が全体に占める割合を示す。
PolypGenのDiceは83.92%、IoUは77.50%で比較法中最高だった。ISIC2017のDiceも83.14%で最高だが、IoUは75.62%でMedSAMの76.55%を下回る。既知のKvasir-SEGではEMCADのDice93.74%に対し92.10%だった。
応用の可能性
編集上の応用案 · 論文が実証した用途とは区別しています。
編集上の応用案:内視鏡画像の確認作業で、医師が撮影した静止画と「ポリープ領域を赤で示す」という指示を入力する。出力は元の組織像を残し、候補領域を赤く重ねた画像とする。
担当者は着色境界を元画像と見比べ、低コントラスト部分や複数病変の見落としを確認する。診断を自動確定する用途ではなく、確認すべき画像を示す作業の補助として検討する。
導入時の検証
導入時の検証案 · 対象データでの再評価が必要です。
導入時の検証案:施設内で利用条件を確認した画像から、撮影条件と病変数が異なる小さな評価群を作る。正解領域を専門家が確認し、InstEditSegと既存の領域分割法へ同じ画像を入力する。
病変ごとのDice、IoU、見落とした病変数を測り、色指定を変えたときの結果も記録する。特に未学習の機器や施設での成績を分けて集計し、担当者が出力を確認する時間も比較する。
限界と課題
著者らは、指定する色への感度があることと、色や形などの属性を文で指定して特定の病変だけを選ぶ機能は未対応だと述べる。したがって、指示文を変えれば任意の病変を自在に選べるとは読めない。
編集上の確認点は、施設や撮影機器が変わった際の境界誤差、色指定の違いによる結果の揺れ、複数病変での取りこぼしである。着色画像から実務用の領域データを得る手順も検証が要る。
出典
元のタイトル・要旨を確認する(英語)
InstEditSeg: Instruction-Driven Image Editing for Polyp and Skin Lesion Segmentation
Accurate segmentation of polyps and skin lesions is pivotal for clinical diagnosis, yet existing methods struggle with low contrast, ambiguous boundaries, and cross-domain distribution discrepancies. Discriminative networks and most diffusion-based segmentation approaches predict standalone binary masks, leaving the visual priors of large-scale pretrained generative models largely unexploited. We propose InstEditSeg, a unified generative framework that reformulates medical segmentation as an instruction-driven image editing problem. Instead of emitting a mask, the model renders a color-coded overlay on the original image, conditioned on a textual instruction, so that the edited output aligns with the natural image distribution learned by latent diffusion models and mitigates the domain gap between natural and medical imagery. To recover fine anatomical structures, we introduce DINOv3 as an auxiliary visual encoder and a DINO Feature Guidance Block that builds a multi-scale feature pyramid. The pyramid is fused into the diffusion U-Net by channel concatenation and zero-initialized convolution so that hierarchical discriminative priors can be injected without perturbing the pretrained weights. A dual-branch classifier-free guidance strategy requiring only two forward passes per denoising step reduces inference cost. On polyp and skin lesion benchmarks the framework achieves accuracy competitive with strong discriminative baselines, and it further demonstrates concrete advantages of the generative formulation: notably better cross-domain generalization on unseen data, more complete multi-lesion segmentation, instruction-conditioned task control, and sampling flexibility. We also analyze the strengths and limitations of the paradigm, including its color sensitivity and unsupported attribute-conditioned selection. Code is available at: https://github.com/wincharm001/InstEditSeg.
論文本体の取得範囲に基づく解説。AIが作成した未校閲の記事です。性能の数値は著者の評価条件に依存します。応用例と検証計画は編集上の提案です。解説更新:2026-09-29