Broadlistening
広聴 AIのBroadlistening パイプラインの Ruby 実装です。LLM を使用して公開コメントをクラスタリング・分析します。
概要
Broadlistening は、大量のコメントや意見を AI を活用して分析するためのパイプラインです。以下のステップで処理を行います:
- Extraction (意見抽出) - コメントから主要な意見を LLM で抽出
- Embedding (ベクトル化) - 抽出した意見をベクトル化
- Clustering (クラスタリング) - UMAP + KMeans + 階層的クラスタリング
- Initial Labelling (初期ラベリング) - 各クラスタに LLM でラベル付け
- Merge Labelling (ラベル統合) - 階層的にラベルを統合
- Overview (概要生成) - 全体の概要を LLM で生成
- Aggregation (JSON 組み立て) - 結果を JSON 形式で出力
インストール
Gemfile に追加
gem 'broadlistening'
または GitHub から直接インストール:
gem 'broadlistening', github: 'takahashim/broadlistening-ruby'
依存関係のインストール
bundle install
使い方
基本的な使用方法
require 'broadlistening'
# コメントデータを準備
comments = [
{ id: "1", body: "環境問題への対策が必要です", proposal_id: "123" },
{ id: "2", body: "公共交通機関の充実を希望します", proposal_id: "123" },
# ...
]
# パイプラインを実行
pipeline = Broadlistening::Pipeline.new(
api_key: ENV['OPENAI_API_KEY'],
model: "gpt-4o-mini",
cluster_nums: [5, 15]
)
result = pipeline.run(comments)
# 結果を取得
puts result[:overview]
puts result[:clusters]
Rails での使用例
# app/jobs/analysis_job.rb
class AnalysisJob < ApplicationJob
queue_as :analysis
def perform(proposal_id)
proposal = Proposal.find(proposal_id)
comments = proposal.comments.map do |c|
{ id: c.id, body: c.body, proposal_id: c.proposal_id }
end
pipeline = Broadlistening::Pipeline.new(
api_key: ENV['OPENAI_API_KEY'],
model: "gpt-4o-mini",
cluster_nums: [5, 15]
)
result = pipeline.run(comments)
proposal.create_analysis_result!(
result_data: result,
comment_count: comments.size
)
end
end
設定オプション
Broadlistening::Pipeline.new(
api_key: "your-api-key", # OpenAI API キー(必須)
model: "gpt-4o-mini", # LLM モデル(デフォルト: gpt-4o-mini)
embedding_model: "text-embedding-3-small", # 埋め込みモデル
cluster_nums: [5, 15], # クラスタ階層の数(デフォルト: [5, 15])
workers: 10, # 並列処理のワーカー数
prompts: { # カスタムプロンプト(オプション)
extraction: "...",
initial_labelling: "...",
merge_labelling: "...",
overview: "..."
}
)
出力形式
パイプラインの結果は以下の構造を持つ Hash です:
{
arguments: [
{
arg_id: "A1_0",
argument: "環境問題への対策が必要",
x: 0.5, # UMAP X座標
y: 0.3, # UMAP Y座標
cluster_ids: ["0", "1_0", "2_3"] # 所属クラスタID
},
# ...
],
clusters: [
{
level: 0,
id: "0",
label: "全体",
description: "",
count: 100,
parent: nil
},
{
level: 1,
id: "1_0",
label: "環境・エネルギー",
description: "環境問題やエネルギー政策に関する意見",
count: 25,
parent: "0"
},
# ...
],
relations: [
{ arg_id: "A1_0", comment_id: "1", proposal_id: "123" },
# ...
],
comment_count: 50,
argument_count: 100,
overview: "分析の概要テキスト...",
config: { model: "gpt-4o-mini", ... }
}
依存関係
- Ruby >= 3.1.0
- activesupport >= 7.0
- numo-narray ~> 0.9
- ruby-openai ~> 7.0
- parallel ~> 1.20
- rice ~> 4.6.0
- umappp ~> 0.2
umappp のインストール
umappp は C++ ネイティブ拡張を含むため、インストール時に C++ コンパイラが必要です:
# macOS
CXX=clang++ gem install umappp
# Linux
gem install umappp
注意: Rice 4.7.x との互換性問題があるため、Rice 4.6.x を使用してください。
開発
# セットアップ
bin/setup
# テスト実行
bundle exec rspec
# コンソール
bin/console
ライセンス
AGPL 3.0