Files
voice_controlled_device/docs/03_hardware_design.md
T

42 KiB
Raw Blame History

ハードウェア・システム詳細設計書 (Voice-Controlled Android App)

1. 使用部品・コンポーネント一覧

1.1 ハードウェア・コンポーネント構成 (Hardware System Layer)

  1. メイン演算・制御ユニット (SoC / Core Processor)
    • 対象端末: Android OS 8.0+ (API Level 26 以上 / Target API Level 34) 搭載の標準 Android 端末 / モジュール (クアッドコア ARM64 2.0GHz 以上推奨)
    • RAM: 3GB 以上
    • ROM/ストレージ: 32GB 以上 (アプリ本体・対話ログ・ローカルモデル保持用)
  2. 音声入力モジュール (Audio Input Device)
    • 内蔵マイクアレイ (Dual/Multi-Microphone Array) または Bluetooth 外部 HFP/A2DP マイク (エレコム LBT-HSC41BK-EC 等)
    • 入力仕様: PCM 16kHz / 16-bit モノラル (Android 音声認識 SpeechRecognizer 前処理入力)
  3. 音声出力モジュール (Audio Output Device)
    • 内蔵スピーカー / D級オーディオアンプ または Bluetooth 外部 A2DP スピーカー / 車載ナビゲーションシステム
    • 出力仕様: PCM 44.1kHz / 48kHz 16-bit ステレオ (Android TextToSpeech 合成出力)
  4. 表示・操作インターフェース (Display & Touch Controller)
    • ディスプレイ: 5インチ〜10インチ静電容量式タッチパネル (1080x1920 解像度等)
    • フィードバック: GPU アクセラレーション対応 (Jetpack Compose 60fps 描画)
  5. 通信モジュール (Network Communications)
    • Wi-Fi (802.11 a/b/g/n/ac) / LTE (4G/5G) セルラー通信モジュール (AI サーバー REST API 通信用)
  6. 外部入力デバイス (Bluetooth Media Control Button)
    • Bluetooth ヘッドセットボタン / リモートメディアキー (AVRCP プロファイル / Intent.ACTION_MEDIA_BUTTON / ACTION_VOICE_COMMAND)

1.2 ソフトウェア・アーキテクチャ概要 (Clean Architecture + MVVM + Foreground Service)

本システムは Clean Architecture, MVVM (Model-View-ViewModel), および常駐 Foreground Service + MediaSessionCompat を組み合わせたアーキテクチャ構造を採用する。レイヤー間の依存関係を一方向 (UI/Service -> Domain <- Voice Engine/Network Layer) に制限し、単体テスト可能性およびバックグラウンド実行時の起動制御・堅牢性を保証する。

  • UI / Presentation Layer: Jetpack Compose によるシングルアクティビティ型 UI、MainViewModel による状態管理と VoiceAppState (Sealed Interface) に基づく単方向データフロー (UDF)。エラー発生時は delay(3000) による 3秒後自動 Idle 復帰制御を実施。
  • Domain Layer: 純粋 Kotlin によるコアビジネスロジック。KanjiToNumberConverter (前処理コンバータ)、ExecuteCommandUseCase、CommandParser、CommandDefinition、CommandId (NAVIGATE_AI 含む)、および各種 ICommandExecutor。
  • Voice Engine Layer: Android プラットフォーム固有 API のラッパーモジュール。IVoiceRecognizer (VoiceRecognitionManager)、ITextToSpeech (TextToSpeechManager)。ITextToSpeech.speak() は Unit を返し StateFlow<TtsStatus> で一元状態管理。
  • Network Layer: Retrofit 2 + OkHttp 4 による AI サーバー REST API クライアント (AiApiService)。認証ヘッダー、タイムアウト設定、Exponential Backoff 自動リトライロジックを保持。
  • Service / Background Layer: Android 14 (API 34) 準拠の常駐フォアグラウンドサービス VoiceAssistantService (foregroundServiceType="microphone|location") および MediaSessionCompat による Bluetooth メディアボタン捕捉・バックグラウンド起動制限 (BAL) 回避モジュール。および AppNotificationListenerService (未読メッセージ取得)。

1.3 ソフトウェアコンポーネント・モジュール詳細

1.3.1 Voice Engine Layer

  • IVoiceRecognizer: android.speech.SpeechRecognizer のラッパー抽象インターフェース。音声入力の開始/停止、認識結果 (Flow/Callback)、マイク音量 (RmsdB) のリアルタイム更新を管理。
  • VoiceRecognitionManager: IVoiceRecognizer の実装クラス。RecognitionListener を内部リスナーとして保持。
  • ITextToSpeech: android.speech.tts.TextToSpeech のラッパー抽象インターフェース。speak(text: String, queueMode: Int = TextToSpeech.QUEUE_FLUSH): Unit のシグネチャを持ち、戻り値型を Unit に統一。
  • TextToSpeechManager: ITextToSpeech の実装クラス。UtteranceProgressListener を用いて Speaking から Idle への状態遷移を管理。

1.3.2 Network Layer (AI API Communication)

  • AiApiService: Retrofit 2 インターフェース。POST /api/v1/navigate を提供。
  • データモデル:
    • AiApiRequest: user_id, query_text, current_location (latitude, longitude), timestamp
    • AiApiResponse: status ("success"), tts_message, destination (name, address, latitude, longitude)
    • AiApiErrorResponse: status ("error"), error_code, message
  • 通信信頼性設計:
    • 認証ヘッダー: Authorization: Bearer <API_KEY> または X-API-Key: <API_KEY>
    • タイムアウト: Connect Timeout 5,000ms / Read Timeout 10,000ms
    • 自動リトライ: Exponential Backoff (初期遅延 1,000ms, 倍率 2.0, 最大 3 回試行)

1.3.3 Domain Layer

  • KanjiToNumberConverter: domain/converter/KanjiToNumberConverter に配置。正規表現パース前の事前処理として「一」「二」「三分」「七時」等の漢数字および全角数字を「1」「2」「3分」「7時」等のアラビア数字に前処理変換。
  • CommandId: コマンド識別子 Enum (GET_TIME_DATE, SET_ALARM_TIMER, CHANGE_SETTINGS, READ_MESSAGES, GET_SYSTEM_INFO, NAVIGATE_AI, UNKNOWN_FALLBACK)。
  • CommandParser: 事前変換後のテキストを入力とし、正規表現マッチングおよびパラメータ抽出を実行。
  • ExecuteCommandUseCase: コマンド解析と実行フローの制御。
  • ICommandExecutor & 実装クラス:
    • GetTimeDateExecutor: システム日時の取得。
    • SetAlarmTimerExecutor: アラーム・タイマーのインテント発行。
    • ChangeSettingsExecutor: 設定変更。
    • ReadMessagesExecutor: NotificationListener 連携による未読メッセージ取得。
    • GetSystemInfoExecutor: バッテリー/通信状態の取得。
    • NavigateAiExecutor: AI サーバー API 通信を行い、目的地の取得後に NaviCon (navicon://point...) URL スキーム生成 (UTF-8 URLエンコード URLEncoder.encode(name, "UTF-8") 適用) および Google Maps / Play Store への段階的フォールバックを実行。
    • UnknownFallbackExecutor: ガイダンス応答。

1.3.4 Service Layer & Background Launch

  • VoiceAssistantService: foregroundServiceType="microphone|location" 指定の常駐 Service。通知バーにインジケータを常駐表示し、バックグラウンドでの音声認識・TTS実行基盤となる。
  • MediaSessionCompat & Callback: Bluetooth メディアボタン (Intent.ACTION_MEDIA_BUTTON / ACTION_VOICE_COMMAND) を MediaSessionCompat の KeyEvent リスナーで直接受託。Android 10+ の Background Activity Launch Restrictions (バックグラウンド起動制限) を回避し、常駐 Service 内から音声認識を即座に開始。
  • AppNotificationListenerService: android.permission.BIND_NOTIFICATION_LISTENER_SERVICE を持つ NotificationListenerService。他アプリの通知本文から未読メッセージを取得。

1.3.5 UI / Presentation Layer

  • MainViewModel: アプリ状態の統合プロセスマネージャー。VoiceAppState.Error 遷移時には viewModelScope 内で delay(3000) を実行し、3秒後に自動で VoiceAppState.Idle へ無害に復帰させる。
  • VoiceAppState: Idle, Listening(val rmsDb: Float), Processing, Speaking(val text: String), Error(val message: String, val errorCode: Int?)
  • Compose Components: MainScreen, StatusBanner, WaveformIndicator, ConversationHistory, ControlBar, PermissionDialog

2. ピン配置・インターフェース定義 (GPIO, SPI, I2C, UART, 電圧レベル等)

2.1 物理・論理インターフェース & バス・信号レベル定義

インターフェース名 物理/論理区分 信号規格・データ形式 物理電圧 / バスレベル 制御・連携方式
マイク音声入力バス (Audio In) 物理 / 論理 PCM 16kHz, 16-bit, Mono 1.8V / 3.3V (ADC / I2S 内蔵バス) AudioRecord / SpeechRecognizer API via Binder IPC
スピーカー音声出力バス (Audio Out) 物理 / 論理 PCM 44.1kHz / 48kHz, 16-bit, Stereo D級アンプ駆動 / DAC 出力 AudioTrack / TextToSpeech API via AudioFlinger
タッチパネル / ディスプレイ (I/O) 物理 MIPI DSI / SPI / I2C (Touch) 1.8V / 3.3V GPIO Android Input subsystem & SurfaceFlinger / Compose GPU
AI サーバー REST API (Network IPC) 論理 HTTPS / JSON (POST /api/v1/navigate) TCP/IP Port 443 (Wi-Fi/Cellular) Retrofit 2 + OkHttp 4 (Auth Header, 5s/10s Timeout, Exp. Backoff)
NaviCon / Google Maps / Play Store IPC 論理 Android Intent / URL Scheme (navicon://point..., geo:..., market://...) OS Binder IPC / Intent Subsystem URLEncoder.encode(name, "UTF-8") 適用 Intent 発行 & パッケージ検出
Bluetooth メディアボタン (AVRCP) 物理 / 論理 Bluetooth HID/AVRCP (ACTION_MEDIA_BUTTON) 2.4GHz RF (Bluetooth HFP/A2DP) MediaSessionCompat.Callback キーハンドリング via VoiceAssistantService
位置情報 API (GPS / Location) 論理 Android Location Manager / FusedLocationProviderClient OS Binder IPC ACCESS_FINE_LOCATION / ACCESS_COARSE_LOCATION
アラーム・タイマー API (AlarmManager) 論理 Android AlarmManager OS Binder IPC SCHEDULE_EXACT_ALARM
通知受託 API (NotificationListener) 論理 NotificationListenerService IPC OS Binder IPC BIND_NOTIFICATION_LISTENER_SERVICE

2.2 クラス構造と抽象インターフェース設計 (Kotlin / Clean Architecture)

2.2.1 クラス図 (Mermaid Diagram)

classDiagram
    namespace UI_Layer {
        class MainScreen {
            +ComposableContent()
        }
        class MainViewModel {
            -ExecuteCommandUseCase executeCommandUseCase
            -IVoiceRecognizer voiceRecognizer
            -ITextToSpeech textToSpeech
            +StateFlow~VoiceAppState~ uiState
            +StateFlow~List~ChatMessage~~ conversationHistory
            +onMicButtonClicked()
            +onPermissionGranted()
            +clearHistory()
        }
        class VoiceAppState {
            <<sealed interface>>
        }
    }

    namespace Domain_Layer {
        class ExecuteCommandUseCase {
            -KanjiToNumberConverter kanjiConverter
            -CommandParser commandParser
            +invoke(String text) Flow~CommandResult~
        }
        class KanjiToNumberConverter {
            +convert(String text) String
        }
        class CommandParser {
            -List~CommandDefinition~ commands
            +parse(String text) CommandMatchResult
        }
        class CommandDefinition {
            +CommandId id
            +List~Regex~ patterns
            +ICommandExecutor executor
        }
        class CommandId {
            <<enum>>
            GET_TIME_DATE
            SET_ALARM_TIMER
            CHANGE_SETTINGS
            READ_MESSAGES
            GET_SYSTEM_INFO
            NAVIGATE_AI
            UNKNOWN_FALLBACK
        }
        class ICommandExecutor {
            <<interface>>
            +execute(CommandMatchResult match) CommandResult
        }
        class NavigateAiExecutor {
            -AiApiService apiService
            -Context context
            +execute(CommandMatchResult match) CommandResult
        }
    }

    namespace Voice_Engine_Layer {
        class IVoiceRecognizer {
            <<interface>>
            +StateFlow~VoiceRecognitionState~ state
            +StateFlow~Float~ rmsDb
            +startListening()
            +stopListening()
            +destroy()
        }
        class VoiceRecognitionManager {
            -SpeechRecognizer speechRecognizer
        }
        class ITextToSpeech {
            <<interface>>
            +StateFlow~TtsStatus~ status
            +speak(String text, Int queueMode) Unit
            +stop()
            +destroy()
        }
        class TextToSpeechManager {
            -TextToSpeech ttsEngine
        }
    }

    namespace Network_Layer {
        class AiApiService {
            <<interface>>
            +navigate(AiApiRequest request) AiApiResponse
        }
        class AiApiRequest {
            +String userId
            +String queryText
            +LocationData currentLocation
            +Long timestamp
        }
        class AiApiResponse {
            +String status
            +String ttsMessage
            +DestinationData destination
        }
    }

    namespace Service_Layer {
        class VoiceAssistantService {
            -MediaSessionCompat mediaSession
            +onStartCommand()
        }
        class AppNotificationListenerService {
            +onNotificationPosted()
        }
    }

    MainScreen ..> MainViewModel : Observe / Event
    MainViewModel --> ExecuteCommandUseCase : Invoke
    MainViewModel --> IVoiceRecognizer : Control
    MainViewModel --> ITextToSpeech : Control
    ExecuteCommandUseCase --> KanjiToNumberConverter : Pre-process
    ExecuteCommandUseCase --> CommandParser : Use
    CommandParser --> CommandDefinition : Contains
    CommandDefinition --> ICommandExecutor : Delegates
    ICommandExecutor <|.. NavigateAiExecutor : Implements
    NavigateAiExecutor --> AiApiService : Call API
    IVoiceRecognizer <|.. VoiceRecognitionManager : Implements
    ITextToSpeech <|.. TextToSpeechManager : Implements
    VoiceAssistantService --> IVoiceRecognizer : Trigger

2.2.2 抽象インターフェース&データモデル定義 (Kotlin コード詳細)

1. Voice Engine Layer インターフェース (speak() 戻り値 Unit 統一)
package com.example.voiceapp.voice

import kotlinx.coroutines.flow.StateFlow

sealed interface VoiceRecognitionState {
    object Idle : VoiceRecognitionState
    object Ready : VoiceRecognitionState
    object Listening : VoiceRecognitionState
    data class Success(val recognizedText: String) : VoiceRecognitionState
    data class Error(val errorCode: Int, val message: String) : VoiceRecognitionState
}

interface IVoiceRecognizer {
    val state: StateFlow<VoiceRecognitionState>
    val rmsDb: StateFlow<Float>

    fun startListening()
    fun stopListening()
    fun destroy()
}

sealed interface TtsStatus {
    object Idle : TtsStatus
    object Speaking : TtsStatus
    object Completed : TtsStatus
    data class Error(val message: String) : TtsStatus
}

interface ITextToSpeech {
    val status: StateFlow<TtsStatus>

    fun speak(text: String, queueMode: Int = 0): Unit
    fun stop()
    fun destroy()
}
2. Network Layer API 定義 & データ構造
package com.example.voiceapp.data.api

import retrofit2.http.Body
import retrofit2.http.Header
import retrofit2.http.POST

data class LocationData(
    val latitude: Double,
    val longitude: Double
)

data class AiApiRequest(
    val user_id: String,
    val query_text: String,
    val current_location: LocationData?,
    val timestamp: Long
)

data class DestinationData(
    val name: String,
    val address: String,
    val latitude: Double,
    val longitude: Double
)

data class AiApiResponse(
    val status: String,
    val tts_message: String,
    val destination: DestinationData?
)

data class AiApiErrorResponse(
    val status: String,
    val error_code: String,
    val message: String
)

interface AiApiService {
    @POST("api/v1/navigate")
    suspend fun navigate(
        @Header("Authorization") authHeader: String,
        @Body request: AiApiRequest
    ): AiApiResponse
}
3. Domain Layer インターフェース、漢数字コンバータ、コマンド構造
package com.example.voiceapp.domain.converter

class KanjiToNumberConverter {
    private val kanjiMap = mapOf(
        '零' to '0', '一' to '1', '二' to '2', '三' to '3', '四' to '4',
        '五' to '5', '六' to '6', '七' to '7', '八' to '8', '九' to '9',
        '0' to '0', '1' to '1', '2' to '2', '3' to '3', '4' to '4',
        '5' to '5', '6' to '6', '7' to '7', '8' to '8', '9' to '9'
    )

    fun convert(input: String): String {
        val sb = StringBuilder()
        for (char in input) {
            val replaced = kanjiMap[char]
            if (replaced != null) {
                sb.append(replaced)
            } else {
                sb.append(char)
            }
        }
        return sb.toString()
    }
}
package com.example.voiceapp.domain.model

enum class CommandId {
    GET_TIME_DATE,
    SET_ALARM_TIMER,
    CHANGE_SETTINGS,
    READ_MESSAGES,
    GET_SYSTEM_INFO,
    NAVIGATE_AI,
    UNKNOWN_FALLBACK
}

data class CommandMatchResult(
    val commandId: CommandId,
    val matchedPattern: String,
    val extractedParameters: Map<String, String>,
    val rawText: String
)

data class CommandResult(
    val commandId: CommandId,
    val isSuccess: Boolean,
    val responseText: String,
    val data: Map<String, Any>? = null
)

interface ICommandExecutor {
    suspend fun execute(matchResult: CommandMatchResult): CommandResult
}

data class CommandDefinition(
    val id: CommandId,
    val patterns: List<Regex>,
    val executor: ICommandExecutor
)
package com.example.voiceapp.domain

import com.example.voiceapp.domain.converter.KanjiToNumberConverter
import com.example.voiceapp.domain.model.*

class CommandParser(private val definitions: List<CommandDefinition>) {
    fun parse(text: String): CommandMatchResult {
        for (def in definitions) {
            for (pattern in def.patterns) {
                val match = pattern.find(text)
                if (match != null) {
                    val params = match.groups.mapIndexedNotNull { index, group ->
                        if (index > 0 && group != null) "param_$index" to group.value else null
                    }.toMap()
                    return CommandMatchResult(def.id, pattern.pattern, params, text)
                }
            }
        }
        return CommandMatchResult(CommandId.UNKNOWN_FALLBACK, "", emptyMap(), text)
    }
}

class ExecuteCommandUseCase(
    private val kanjiConverter: KanjiToNumberConverter,
    private val commandParser: CommandParser
) {
    suspend operator fun invoke(inputText: String): CommandResult {
        val normalizedText = kanjiConverter.convert(inputText)
        val matchResult = commandParser.parse(normalizedText)
        return matchResult.executor.execute(matchResult)
    }
}
4. Presentation Layer 状態・メッセージモデル
package com.example.voiceapp.ui

import com.example.voiceapp.domain.model.CommandId

sealed interface VoiceAppState {
    object Idle : VoiceAppState
    data class Listening(val rmsDb: Float = 0f) : VoiceAppState
    object Processing : VoiceAppState
    data class Speaking(val text: String) : VoiceAppState
    data class Error(val message: String, val errorCode: Int? = null) : VoiceAppState
}

enum class MessageSender { USER, SYSTEM }

data class ChatMessage(
    val id: String = java.util.UUID.randomUUID().toString(),
    val sender: MessageSender,
    val text: String,
    val timestamp: Long = System.currentTimeMillis(),
    val commandId: CommandId? = null
)

2.2.3 依存性注入 (DI: Hilt Framework Design)

package com.example.voiceapp.di

import android.content.Context
import com.example.voiceapp.data.api.AiApiService
import com.example.voiceapp.domain.*
import com.example.voiceapp.domain.converter.KanjiToNumberConverter
import com.example.voiceapp.domain.executor.*
import com.example.voiceapp.domain.model.*
import com.example.voiceapp.voice.*
import dagger.Module
import dagger.Provides
import dagger.hilt.InstallIn
import dagger.hilt.android.qualifiers.ApplicationContext
import dagger.hilt.components.SingletonComponent
import okhttp3.OkHttpClient
import retrofit2.Retrofit
import retrofit2.converter.gson.GsonConverterFactory
import java.util.concurrent.TimeUnit
import javax.inject.Singleton

@Module
@InstallIn(SingletonComponent::class)
object VoiceModule {

    @Provides
    @Singleton
    fun provideVoiceRecognizer(
        @ApplicationContext context: Context
    ): IVoiceRecognizer = VoiceRecognitionManager(context)

    @Provides
    @Singleton
    fun provideTextToSpeech(
        @ApplicationContext context: Context
    ): ITextToSpeech = TextToSpeechManager(context)
}

@Module
@InstallIn(SingletonComponent::class)
object NetworkModule {

    @Provides
    @Singleton
    fun provideOkHttpClient(): OkHttpClient {
        return OkHttpClient.Builder()
            .connectTimeout(5000, TimeUnit.MILLISECONDS)
            .readTimeout(10000, TimeUnit.MILLISECONDS)
            .addInterceptor { chain ->
                var request = chain.request()
                val apiKey = "YOUR_API_KEY_HERE"
                request = request.newBuilder()
                    .header("Authorization", "Bearer $apiKey")
                    .build()
                
                // Exponential Backoff Retry (Max 3 Tries)
                var response = chain.proceed(request)
                var tryCount = 0
                var backoffDelay = 1000L
                while (!response.isSuccessful && tryCount < 3) {
                    tryCount++
                    Thread.sleep(backoffDelay)
                    backoffDelay *= 2
                    response.close()
                    response = chain.proceed(request)
                }
                response
            }
            .build()
    }

    @Provides
    @Singleton
    fun provideAiApiService(okHttpClient: OkHttpClient): AiApiService {
        return Retrofit.Builder()
            .baseUrl("https://api.example.com/")
            .client(okHttpClient)
            .addConverterFactory(GsonConverterFactory.create())
            .build()
            .create(AiApiService::class.java)
    }
}

@Module
@InstallIn(SingletonComponent::class)
object DomainModule {

    @Provides
    @Singleton
    fun provideKanjiToNumberConverter(): KanjiToNumberConverter {
        return KanjiToNumberConverter()
    }

    @Provides
    @Singleton
    fun provideNavigateAiExecutor(
        @ApplicationContext context: Context,
        apiService: AiApiService
    ): NavigateAiExecutor {
        return NavigateAiExecutor(context, apiService)
    }

    @Provides
    @Singleton
    fun provideCommandDefinitions(
        @ApplicationContext context: Context,
        navigateAiExecutor: NavigateAiExecutor
    ): List<CommandDefinition> {
        return listOf(
            CommandDefinition(
                id = CommandId.NAVIGATE_AI,
                patterns = listOf(Regex(".*(ナビ|行きたい|向かう|セット).*")),
                executor = navigateAiExecutor
            ),
            CommandDefinition(
                id = CommandId.GET_TIME_DATE,
                patterns = listOf(Regex(".*(今何時|何時ですか|日付|今日は何日).*")),
                executor = GetTimeDateExecutor()
            ),
            CommandDefinition(
                id = CommandId.SET_ALARM_TIMER,
                patterns = listOf(Regex(".*(\\d+)\\s*(分|秒|時間)タイマー.*|.*(\\d+)時にアラーム.*")),
                executor = SetAlarmTimerExecutor(context)
            ),
            CommandDefinition(
                id = CommandId.CHANGE_SETTINGS,
                patterns = listOf(Regex(".*(音量を|ダークモード|ライトモード).*")),
                executor = ChangeSettingsExecutor(context)
            ),
            CommandDefinition(
                id = CommandId.READ_MESSAGES,
                patterns = listOf(Regex(".*(メッセージ|未読|通知).*")),
                executor = ReadMessagesExecutor(context)
            ),
            CommandDefinition(
                id = CommandId.GET_SYSTEM_INFO,
                patterns = listOf(Regex(".*(バッテリー|電池|Wi-Fi|ネットワーク).*")),
                executor = GetSystemInfoExecutor(context)
            ),
            CommandDefinition(
                id = CommandId.UNKNOWN_FALLBACK,
                patterns = emptyList(),
                executor = UnknownFallbackExecutor()
            )
        )
    }

    @Provides
    @Singleton
    fun provideCommandParser(definitions: List<CommandDefinition>): CommandParser {
        return CommandParser(definitions)
    }

    @Provides
    @Singleton
    fun provideExecuteCommandUseCase(
        kanjiConverter: KanjiToNumberConverter,
        parser: CommandParser
    ): ExecuteCommandUseCase {
        return ExecuteCommandUseCase(kanjiConverter, parser)
    }
}

3. 回路・構造・システムの注意事項

3.1 物理・ハードウェア、電力、オーディオ、ネットワーク注意事項

  1. マイクフィードバック (AEC) & オーディオフォーカス制御

    • TTS 発話中のスピーカー音量をマイクが拾うループを完全防止するため、ITextToSpeech の status が Speaking の間は IVoiceRecognizer の startListening() 呼び出しをシステムレベルでロックする。
    • speak() 呼出時に AudioManager.requestAudioFocus(AUDIOFOCUS_GAIN_TRANSIENT_MAY_DUCK) を要求し、他アプリの音量を自動ミュート/減衰させる。発話完了 (onDone) 時に abandonAudioFocus() を実行する。
  2. バッテリー消費抑制 & Bluetooth バックグラウンド起動制御 (BAL 対策)

    • Android 10+ (API 29+) のスリープ中・バックグラウンドからの Activity 起動制限 (Background Activity Launch restrictions) に対応するため、VoiceAssistantService を常駐 Foreground Service (foregroundServiceType="microphone|location") として動作させる。
    • Bluetooth ヘッドセットボタン (Intent.ACTION_MEDIA_BUTTON) 押下時は MediaSessionCompat の KeyEventListener でイベントを直接受信し、Service 内でマイク録音・音声認識を開始する。不要な Activity ダイレクト起動を行わないことで制限を回避し、かつ常時マイク監視を行わないため CPU の省電力サスペンドを維持する。
  3. 通信信頼性・エラーハンドリング (Exponential Backoff & タイムアウト)

    • モバイル回線接続の不安定性に備え、REST API 通信には Connect Timeout 5,000ms, Read Timeout 10,000ms を設定。
    • API 通信失敗時は Exponential Backoff (初期遅延 1,000ms, 倍率 2, 最大 3 回) にて自動リトライ。失敗継続時は AiApiErrorResponse (LOCATION_NOT_FOUND 等) を受託し、ユーザーへ親切なガイド音声を読み上げる。

3.2 システム制御シーケンス (シーケンス図)

3.2.1 AI ナビゲーション実行 & 段階的フォールバックシーケンス

sequenceDiagram
    autonumber
    actor User as ユーザー
    participant UI as Compose UI (MainScreen)
    participant VM as MainViewModel
    participant VR as IVoiceRecognizer
    participant UC as ExecuteCommandUseCase
    participant KC as KanjiToNumberConverter
    participant CP as CommandParser
    participant EX as NavigateAiExecutor
    participant API as AiApiService (REST API)
    participant TTS as ITextToSpeech

    User->>UI: マイクボタンタップ / Bluetoothボタン
    UI->>VM: onMicButtonClicked()
    VM->>VR: startListening()
    VR->>VM: state = Listening(rmsDb)
    VM->>UI: StateFlow Update: VoiceAppState.Listening

    User->>VR: 発話 (例:「仕事で行く三分後の〇〇会社に向かう」)
    VR->>VM: state = Success("仕事で行く三分後の〇〇会社に向かう")
    VM->>UI: StateFlow Update: VoiceAppState.Processing
    VM->>VM: 会話履歴追加 (USER)

    VM->>UC: invoke("仕事で行く三分後の〇〇会社に向かう")
    UC->>KC: convert("仕事で行く三分後の〇〇会社に向かう")
    KC-->>UC: 前処理結果 ("仕事で行く3分後の〇〇会社に向かう")
    UC->>CP: parse("仕事で行く3分後の〇〇会社に向かう")
    CP-->>UC: CommandMatchResult (NAVIGATE_AI)
    UC->>EX: NavigateAiExecutor.execute()
    
    EX->>API: POST /api/v1/navigate (Auth Header, 5s/10s Timeout)
    API-->>EX: HTTP 200 (AiApiResponse: tts_message, destination)
    
    EX->>EX: UTF-8 URL Encode: URLEncoder.encode(name, "UTF-8")
    alt NaviCon インストール検出時
        EX->>EX: NaviCon Intent (navicon://point?ll=lat,lng&title=encoded_name)
    else NaviCon 未検出 且つ Google Maps インストール時
        EX->>EX: Google Maps Intent (geo:lat,lng?q=encoded_name)
    else いずれも未検出時
        EX->>EX: Play Store Intent (market://details?id=jp.co.denso.navicon.user)
    end

    EX-->>UC: CommandResult (responseText = tts_message)
    UC-->>VM: CommandResult 返却

    VM->>TTS: speak(tts_message)
    TTS->>VM: status = Speaking
    VM->>UI: StateFlow Update: VoiceAppState.Speaking
    TTS->>User: 音声出力
    TTS->>VM: status = Completed
    VM->>UI: StateFlow Update: VoiceAppState.Idle

3.2.2 エラー発生時 3秒自動復帰シーケンス (delay 3000ms)

sequenceDiagram
    autonumber
    actor User as ユーザー
    participant UI as MainScreen
    participant VM as MainViewModel
    participant VR as IVoiceRecognizer
    participant TTS as ITextToSpeech

    User->>VR: 発話不能 / ノイズ入力
    VR->>VM: state = Error(ERROR_SPEECH_TIMEOUT)
    VM->>UI: StateFlow Update: VoiceAppState.Error("音声が認識できませんでした")
    VM->>TTS: speak("音声が認識できませんでした")
    
    note over VM: MainViewModel 内で Coroutine delay(3000) 起動
    VM->>VM: viewModelScope.launch { delay(3000); updateState(VoiceAppState.Idle) }
    
    note over VM,UI: 3秒間エラーメッセージ・ダイアログを表示
    VM->>UI: StateFlow Update: VoiceAppState.Idle
    UI->>User: Idle 状態表示 (自動安全復帰完了)

3.2.3 Bluetooth メディアボタン バックグラウンド起動シーケンス

sequenceDiagram
    autonumber
    actor User as ユーザー (ヘッドセット)
    participant BT as Bluetooth Headset
    participant SVC as VoiceAssistantService (Foreground)
    participant MS as MediaSessionCompat.Callback
    participant VR as IVoiceRecognizer
    participant VM as MainViewModel

    User->>BT: メディアボタン押下
    BT->>MS: Intent.ACTION_MEDIA_BUTTON
    MS->>SVC: onMediaButtonEvent()
    SVC->>SVC: チェック: 常駐 Foreground Service 動作中
    SVC->>VR: startListening()
    VR->>VM: state = Listening
    VM->>User: 音声認識開始 (画面オフ状態でも制御継続)

3.3 ナビゲーション連携 (NaviCon & Google Maps) & URL エンコード仕様

  1. NaviCon 連携 (メイン軸 / 完全無料)
    • API 応答から受信した目的地名 (name) に必ず URLEncoder.encode(name, "UTF-8") 処理を行う。
    • スキームフォーマット: navicon://point?ll={latitude},{longitude}&title={encoded_name}
  2. Google Maps 連携 (フォールバック 1)
    • スキームフォーマット: geo:{latitude},{longitude}?q={encoded_name}
  3. Google Play ストア誘導 (フォールバック 2)
    • スキームフォーマット: market://details?id=jp.co.denso.navicon.user
  4. パッケージ確認ロジック (フォールバック判定順序)
    • PackageManager.getPackageInfo("jp.co.denso.navicon.user", 0) を試行。
    • 存在する場合 -> NaviCon Intent 発行。
    • 存在しない場合 -> 「NaviConが未検出のため、Google Mapsで表示します」と TTS 発話後、Google Maps Intent 発行。
    • Google Maps も未検出の場合 -> Playストア Intent 発行。

3.4 ディレクトリ構造とソースファイル配置設計

ルートディレクトリ構成からアプリ内部のパッケージ構成に至る完全なディレクトリ構造は以下の通り。

.
├── .gitea/
│   └── workflows/
│       └── build.yaml                  # Gitea Runner CI/CD ワークフロー
├── docker/
│   ├── Dockerfile                      # Android ビルド用 Dockerfile (Ubuntu 22.04 + JDK 17 + Android SDK)
│   └── entrypoint.sh                   # コンテナ起動初期化スクリプト
├── docker-compose.yml                  # ローカル Docker ビルド定義
├── app/
│   ├── build.gradle.kts                # アプリ依存関係設定
│   └── src/
│       ├── main/
│       │   ├── java/com/example/voiceapp/
│       │   │   ├── VoiceApplication.kt # Application クラス (@HiltAndroidApp)
│       │   │   ├── MainActivity.kt     # シングルアクティビティ
│       │   │   ├── data/
│       │   │   │   └── api/
│       │   │   │       ├── AiApiService.kt
│       │   │   │       └── models/     # AiApiRequest, AiApiResponse, AiApiErrorResponse
│       │   │   ├── di/
│       │   │   │   ├── VoiceModule.kt
│       │   │   │   ├── NetworkModule.kt
│       │   │   │   └── DomainModule.kt
│       │   │   ├── domain/
│       │   │   │   ├── converter/
│       │   │   │   │   └── KanjiToNumberConverter.kt
│       │   │   │   ├── executor/
│       │   │   │   │   ├── ICommandExecutor.kt
│       │   │   │   │   ├── GetTimeDateExecutor.kt
│       │   │   │   │   ├── SetAlarmTimerExecutor.kt
│       │   │   │   │   ├── ChangeSettingsExecutor.kt
│       │   │   │   │   ├── ReadMessagesExecutor.kt
│       │   │   │   │   ├── GetSystemInfoExecutor.kt
│       │   │   │   │   ├── NavigateAiExecutor.kt
│       │   │   │   │   └── UnknownFallbackExecutor.kt
│       │   │   │   ├── model/
│       │   │   │   │   ├── CommandId.kt
│       │   │   │   │   ├── CommandMatchResult.kt
│       │   │   │   │   └── CommandResult.kt
│       │   │   │   ├── CommandDefinition.kt
│       │   │   │   ├── CommandParser.kt
│       │   │   │   └── ExecuteCommandUseCase.kt
│       │   │   ├── voice/
│       │   │   │   ├── IVoiceRecognizer.kt
│       │   │   │   ├── VoiceRecognitionManager.kt
│       │   │   │   ├── ITextToSpeech.kt
│       │   │   │   └── TextToSpeechManager.kt
│       │   │   ├── service/
│       │   │   │   ├── VoiceAssistantService.kt
│       │   │   │   └── AppNotificationListenerService.kt
│       │   │   └── ui/
│       │   │       ├── VoiceAppState.kt
│       │   │       ├── ChatMessage.kt
│       │   │       ├── MainViewModel.kt
│       │   │       └── components/
│       │   │           ├── MainScreen.kt
│       │   │           ├── StatusBanner.kt
│       │   │           ├── WaveformIndicator.kt
│       │   │           ├── ConversationHistory.kt
│       │   │           ├── ControlBar.kt
│       │   │           └── PermissionDialog.kt
│       │   └── AndroidManifest.xml
│       └── test/
│           └── java/com/example/voiceapp/
│               ├── domain/
│               │   ├── KanjiToNumberConverterTest.kt
│               │   ├── CommandParserTest.kt
│               │   └── ExecuteCommandUseCaseTest.kt
│               └── ui/
│                   └── MainViewModelTest.kt
└── README.md

3.5 AndroidManifest.xml 権限設定・サービス設定設計

Android OS 8.0+ (API Level 26) から Android 14 (API Level 34) までの互換性および必須パーミッション・フォアグラウンドサービス宣言を網羅する AndroidManifest.xml の完全記述。

<?xml version="1.0" encoding="utf-8"?>
<manifest xmlns:android="http://schemas.android.com/apk/res/android"
    package="com.example.voiceapp">

    <!-- 必須パーミッション宣言 -->
    <!-- 1. 音声入力権限 (危険権限) -->
    <uses-permission android:name="android.permission.RECORD_AUDIO" />

    <!-- 2. AI サーバー通信・ネットワーク確認権限 -->
    <uses-permission android:name="android.permission.INTERNET" />
    <uses-permission android:name="android.permission.ACCESS_NETWORK_STATE" />

    <!-- 3. 位置情報権限 (AI ナビゲーション現在地付与用) -->
    <uses-permission android:name="android.permission.ACCESS_FINE_LOCATION" />
    <uses-permission android:name="android.permission.ACCESS_COARSE_LOCATION" />

    <!-- 4. オーディオ設定(音量変更・フォーカス要求)権限 -->
    <uses-permission android:name="android.permission.MODIFY_AUDIO_SETTINGS" />

    <!-- 5. アラーム・タイマーインテント設定権限 (Android 12+ API 31+) -->
    <uses-permission android:name="android.permission.SCHEDULE_EXACT_ALARM" />

    <!-- 6. 通知出力権限 (Android 13+ API 33+) -->
    <uses-permission android:name="android.permission.POST_NOTIFICATIONS" />

    <!-- 7. 未読メッセージ通知受託権限 -->
    <uses-permission android:name="android.permission.BIND_NOTIFICATION_LISTENER_SERVICE" />

    <!-- 8. Android 14 (API 34) 必須フォアグラウンドサービス権限 -->
    <uses-permission android:name="android.permission.FOREGROUND_SERVICE" />
    <uses-permission android:name="android.permission.FOREGROUND_SERVICE_MICROPHONE" />
    <uses-permission android:name="android.permission.FOREGROUND_SERVICE_LOCATION" />

    <!-- ハードウェア機能制限の設定 (マイク必須設定) -->
    <uses-feature
        android:name="android.hardware.microphone"
        android:required="true" />
    <uses-feature
        android:name="android.hardware.location.gps"
        android:required="false" />

    <application
        android:name=".VoiceApplication"
        android:allowBackup="true"
        android:icon="@mipmap/ic_launcher"
        android:label="音声操作ナビ端末"
        android:roundIcon="@mipmap/ic_launcher_round"
        android:supportsRtl="true"
        android:theme="@style/Theme.VoiceControlledDevice">

        <!-- シングルアクティビティ構成 -->
        <activity
            android:name=".MainActivity"
            android:exported="true"
            android:launchMode="singleTop"
            android:screenOrientation="portrait"
            android:windowSoftInputMode="adjustResize">
            <intent-filter>
                <action android:name="android.intent.action.MAIN" />
                <category android:name="android.intent.category.LAUNCHER" />
            </intent-filter>

            <!-- 音声アシスタント呼び出しインテント -->
            <intent-filter>
                <action android:name="android.intent.action.VOICE_COMMAND" />
                <category android:name="android.intent.category.DEFAULT" />
            </intent-filter>
        </activity>

        <!-- 常駐フォアグラウンドサービス (Android 14 API 34 対応) -->
        <service
            android:name=".service.VoiceAssistantService"
            android:exported="false"
            android:foregroundServiceType="microphone|location" />

        <!-- 未読メッセージ読み上げ通知受託サービス -->
        <service
            android:name=".service.AppNotificationListenerService"
            android:label="メッセージ読み上げサービス"
            android:permission="android.permission.BIND_NOTIFICATION_LISTENER_SERVICE"
            android:exported="true">
            <intent-filter>
                <action android:name="android.service.notification.NotificationListenerService" />
            </intent-filter>
        </service>

        <!-- SpeechRecognizer, NaviCon, Google Maps パッケージクエリ設定 (Android 11+ API 30+) -->
        <queries>
            <intent>
                <action android:name="android.speech.RecognitionService" />
            </intent>
            <intent>
                <action android:name="android.intent.action.TTS_SERVICE" />
            </intent>
            <package android:name="jp.co.denso.navicon.user" />
            <package android:name="com.google.android.apps.maps" />
        </queries>

    </application>

</manifest>

3.6 Gitea CI/CD Docker Runner 環境・ビルドパイプライン設計

  1. Docker Runner ビルドモデル

    • CI/CD パイプラインは Gitea Act Runner の Docker Runner モードにて動的コンテナ環境を生成して実行する。
    • ホストOSに依存せず、コンテナイメージ (docker/Dockerfile) 内で JDK 17, Android SDK (API 34 Build-tools), Gradle 8.x 環境を完全にカプセル化する。
  2. Gitea ワークフロー定義 (.gitea/workflows/build.yaml)

    name: Android CI Build & Test
    
    on:
      push:
        branches: [ main, develop ]
      pull_request:
        branches: [ main, develop ]
    
    jobs:
      build-and-test:
        runs-on: ubuntu-latest
        container:
          image: voice-app-build:latest
        steps:
          - name: Checkout Repository
            uses: actions/checkout@v3
    
          - name: Run Unit Tests
            run: ./gradlew testDebugUnitTest --stacktrace
    
          - name: Assemble Debug APK
            run: ./gradlew assembleDebug --stacktrace
    
          - name: Upload Artifacts
            uses: actions/upload-artifact@v3
            with:
              name: debug-apk
              path: app/build/outputs/apk/debug/app-debug.apk
    
  3. ローカル検証モデル (docker-compose.yml)

    • ローカル環境でも docker-compose run --rm build-env ./gradlew testDebugUnitTest で CI と同一のテスト・ビルド環境を実行可能とする。