Initial commit: Voice controlled device application, specs and Gitea Actions workflow

This commit is contained in:
2026-08-10 08:26:00 +00:00
commit d449469116
69 changed files with 3209 additions and 0 deletions
+957
View File
@@ -0,0 +1,957 @@
# ハードウェア・システム詳細設計書 (Voice-Controlled Android App)
## 1. 使用部品・コンポーネント一覧
### 1.1 ハードウェア・コンポーネント構成 (Hardware System Layer)
1. **メイン演算・制御ユニット (SoC / Core Processor)**
- 対象端末: Android OS 8.0+ (API Level 26 以上 / Target API Level 34) 搭載の標準 Android 端末 / モジュール (クアッドコア ARM64 2.0GHz 以上推奨)
- RAM: 3GB 以上
- ROM/ストレージ: 32GB 以上 (アプリ本体・対話ログ・ローカルモデル保持用)
2. **音声入力モジュール (Audio Input Device)**
- 内蔵マイクアレイ (Dual/Multi-Microphone Array) または Bluetooth 外部 HFP/A2DP マイク (エレコム LBT-HSC41BK-EC 等)
- 入力仕様: PCM 16kHz / 16-bit モノラル (Android 音声認識 `SpeechRecognizer` 前処理入力)
3. **音声出力モジュール (Audio Output Device)**
- 内蔵スピーカー / D級オーディオアンプ または Bluetooth 外部 A2DP スピーカー / 車載ナビゲーションシステム
- 出力仕様: PCM 44.1kHz / 48kHz 16-bit ステレオ (Android `TextToSpeech` 合成出力)
4. **表示・操作インターフェース (Display & Touch Controller)**
- ディスプレイ: 5インチ〜10インチ静電容量式タッチパネル (1080x1920 解像度等)
- フィードバック: GPU アクセラレーション対応 (Jetpack Compose 60fps 描画)
5. **通信モジュール (Network Communications)**
- Wi-Fi (802.11 a/b/g/n/ac) / LTE (4G/5G) セルラー通信モジュール (AI サーバー REST API 通信用)
6. **外部入力デバイス (Bluetooth Media Control Button)**
- Bluetooth ヘッドセットボタン / リモートメディアキー (AVRCP プロファイル / `Intent.ACTION_MEDIA_BUTTON` / `ACTION_VOICE_COMMAND`)
### 1.2 ソフトウェア・アーキテクチャ概要 (Clean Architecture + MVVM + Foreground Service)
本システムは Clean Architecture, MVVM (Model-View-ViewModel), および常駐 Foreground Service + `MediaSessionCompat` を組み合わせたアーキテクチャ構造を採用する。レイヤー間の依存関係を一方向 (UI/Service -> Domain <- Voice Engine/Network Layer) に制限し、単体テスト可能性およびバックグラウンド実行時の起動制御・堅牢性を保証する。
- **UI / Presentation Layer**: Jetpack Compose によるシングルアクティビティ型 UI、`MainViewModel` による状態管理と `VoiceAppState` (Sealed Interface) に基づく単方向データフロー (UDF)。エラー発生時は `delay(3000)` による 3秒後自動 Idle 復帰制御を実施。
- **Domain Layer**: 純粋 Kotlin によるコアビジネスロジック。`KanjiToNumberConverter` (前処理コンバータ)、`ExecuteCommandUseCase`、`CommandParser`、`CommandDefinition`、`CommandId` (`NAVIGATE_AI` 含む)、および各種 `ICommandExecutor`。
- **Voice Engine Layer**: Android プラットフォーム固有 API のラッパーモジュール。`IVoiceRecognizer` (`VoiceRecognitionManager`)、`ITextToSpeech` (`TextToSpeechManager`)。`ITextToSpeech.speak()` は `Unit` を返し `StateFlow<TtsStatus>` で一元状態管理。
- **Network Layer**: Retrofit 2 + OkHttp 4 による AI サーバー REST API クライアント (`AiApiService`)。認証ヘッダー、タイムアウト設定、Exponential Backoff 自動リトライロジックを保持。
- **Service / Background Layer**: Android 14 (API 34) 準拠の常駐フォアグラウンドサービス `VoiceAssistantService` (`foregroundServiceType="microphone|location"`) および `MediaSessionCompat` による Bluetooth メディアボタン捕捉・バックグラウンド起動制限 (BAL) 回避モジュール。および `AppNotificationListenerService` (未読メッセージ取得)。
### 1.3 ソフトウェアコンポーネント・モジュール詳細
#### 1.3.1 Voice Engine Layer
- **`IVoiceRecognizer`**: `android.speech.SpeechRecognizer` のラッパー抽象インターフェース。音声入力の開始/停止、認識結果 (Flow/Callback)、マイク音量 (RmsdB) のリアルタイム更新を管理。
- **`VoiceRecognitionManager`**: `IVoiceRecognizer` の実装クラス。`RecognitionListener` を内部リスナーとして保持。
- **`ITextToSpeech`**: `android.speech.tts.TextToSpeech` のラッパー抽象インターフェース。`speak(text: String, queueMode: Int = TextToSpeech.QUEUE_FLUSH): Unit` のシグネチャを持ち、戻り値型を `Unit` に統一。
- **`TextToSpeechManager`**: `ITextToSpeech` の実装クラス。`UtteranceProgressListener` を用いて `Speaking` から `Idle` への状態遷移を管理。
#### 1.3.2 Network Layer (AI API Communication)
- **`AiApiService`**: Retrofit 2 インターフェース。POST `/api/v1/navigate` を提供。
- **データモデル**:
- `AiApiRequest`: `user_id`, `query_text`, `current_location` (latitude, longitude), `timestamp`
- `AiApiResponse`: `status` ("success"), `tts_message`, `destination` (name, address, latitude, longitude)
- `AiApiErrorResponse`: `status` ("error"), `error_code`, `message`
- **通信信頼性設計**:
- 認証ヘッダー: `Authorization: Bearer <API_KEY>` または `X-API-Key: <API_KEY>`
- タイムアウト: Connect Timeout 5,000ms / Read Timeout 10,000ms
- 自動リトライ: Exponential Backoff (初期遅延 1,000ms, 倍率 2.0, 最大 3 回試行)
#### 1.3.3 Domain Layer
- **`KanjiToNumberConverter`**: `domain/converter/KanjiToNumberConverter` に配置。正規表現パース前の事前処理として「一」「二」「三分」「七時」等の漢数字および全角数字を「1」「2」「3分」「7時」等のアラビア数字に前処理変換。
- **`CommandId`**: コマンド識別子 Enum (`GET_TIME_DATE`, `SET_ALARM_TIMER`, `CHANGE_SETTINGS`, `READ_MESSAGES`, `GET_SYSTEM_INFO`, `NAVIGATE_AI`, `UNKNOWN_FALLBACK`)。
- **`CommandParser`**: 事前変換後のテキストを入力とし、正規表現マッチングおよびパラメータ抽出を実行。
- **`ExecuteCommandUseCase`**: コマンド解析と実行フローの制御。
- **`ICommandExecutor` & 実装クラス**:
- `GetTimeDateExecutor`: システム日時の取得。
- `SetAlarmTimerExecutor`: アラーム・タイマーのインテント発行。
- `ChangeSettingsExecutor`: 設定変更。
- `ReadMessagesExecutor`: NotificationListener 連携による未読メッセージ取得。
- `GetSystemInfoExecutor`: バッテリー/通信状態の取得。
- `NavigateAiExecutor`: AI サーバー API 通信を行い、目的地の取得後に NaviCon (`navicon://point...`) URL スキーム生成 (UTF-8 URLエンコード `URLEncoder.encode(name, "UTF-8")` 適用) および Google Maps / Play Store への段階的フォールバックを実行。
- `UnknownFallbackExecutor`: ガイダンス応答。
#### 1.3.4 Service Layer & Background Launch
- **`VoiceAssistantService`**: `foregroundServiceType="microphone|location"` 指定の常駐 Service。通知バーにインジケータを常駐表示し、バックグラウンドでの音声認識・TTS実行基盤となる。
- **`MediaSessionCompat` & Callback**: Bluetooth メディアボタン (`Intent.ACTION_MEDIA_BUTTON` / `ACTION_VOICE_COMMAND`) を `MediaSessionCompat` の KeyEvent リスナーで直接受託。Android 10+ の Background Activity Launch Restrictions (バックグラウンド起動制限) を回避し、常駐 Service 内から音声認識を即座に開始。
- **`AppNotificationListenerService`**: `android.permission.BIND_NOTIFICATION_LISTENER_SERVICE` を持つ NotificationListenerService。他アプリの通知本文から未読メッセージを取得。
#### 1.3.5 UI / Presentation Layer
- **`MainViewModel`**: アプリ状態の統合プロセスマネージャー。`VoiceAppState.Error` 遷移時には `viewModelScope` 内で `delay(3000)` を実行し、3秒後に自動で `VoiceAppState.Idle` へ無害に復帰させる。
- **`VoiceAppState`**: `Idle`, `Listening(val rmsDb: Float)`, `Processing`, `Speaking(val text: String)`, `Error(val message: String, val errorCode: Int?)`
- **Compose Components**: `MainScreen`, `StatusBanner`, `WaveformIndicator`, `ConversationHistory`, `ControlBar`, `PermissionDialog`
## 2. ピン配置・インターフェース定義 (GPIO, SPI, I2C, UART, 電圧レベル等)
### 2.1 物理・論理インターフェース & バス・信号レベル定義
| インターフェース名 | 物理/論理区分 | 信号規格・データ形式 | 物理電圧 / バスレベル | 制御・連携方式 |
|---|---|---|---|---|
| **マイク音声入力バス (Audio In)** | 物理 / 論理 | PCM 16kHz, 16-bit, Mono | 1.8V / 3.3V (ADC / I2S 内蔵バス) | `AudioRecord` / `SpeechRecognizer` API via Binder IPC |
| **スピーカー音声出力バス (Audio Out)** | 物理 / 論理 | PCM 44.1kHz / 48kHz, 16-bit, Stereo | D級アンプ駆動 / DAC 出力 | `AudioTrack` / `TextToSpeech` API via AudioFlinger |
| **タッチパネル / ディスプレイ (I/O)** | 物理 | MIPI DSI / SPI / I2C (Touch) | 1.8V / 3.3V GPIO | Android Input subsystem & SurfaceFlinger / Compose GPU |
| **AI サーバー REST API (Network IPC)** | 論理 | HTTPS / JSON (POST /api/v1/navigate) | TCP/IP Port 443 (Wi-Fi/Cellular) | Retrofit 2 + OkHttp 4 (Auth Header, 5s/10s Timeout, Exp. Backoff) |
| **NaviCon / Google Maps / Play Store IPC** | 論理 | Android Intent / URL Scheme (`navicon://point...`, `geo:...`, `market://...`) | OS Binder IPC / Intent Subsystem | `URLEncoder.encode(name, "UTF-8")` 適用 Intent 発行 & パッケージ検出 |
| **Bluetooth メディアボタン (AVRCP)** | 物理 / 論理 | Bluetooth HID/AVRCP (`ACTION_MEDIA_BUTTON`) | 2.4GHz RF (Bluetooth HFP/A2DP) | `MediaSessionCompat.Callback` キーハンドリング via `VoiceAssistantService` |
| **位置情報 API (GPS / Location)** | 論理 | Android Location Manager / FusedLocationProviderClient | OS Binder IPC | `ACCESS_FINE_LOCATION` / `ACCESS_COARSE_LOCATION` |
| **アラーム・タイマー API (AlarmManager)** | 論理 | Android AlarmManager | OS Binder IPC | `SCHEDULE_EXACT_ALARM` |
| **通知受託 API (NotificationListener)** | 論理 | NotificationListenerService IPC | OS Binder IPC | `BIND_NOTIFICATION_LISTENER_SERVICE` |
### 2.2 クラス構造と抽象インターフェース設計 (Kotlin / Clean Architecture)
#### 2.2.1 クラス図 (Mermaid Diagram)
```mermaid
classDiagram
namespace UI_Layer {
class MainScreen {
+ComposableContent()
}
class MainViewModel {
-ExecuteCommandUseCase executeCommandUseCase
-IVoiceRecognizer voiceRecognizer
-ITextToSpeech textToSpeech
+StateFlow~VoiceAppState~ uiState
+StateFlow~List~ChatMessage~~ conversationHistory
+onMicButtonClicked()
+onPermissionGranted()
+clearHistory()
}
class VoiceAppState {
<<sealed interface>>
}
}
namespace Domain_Layer {
class ExecuteCommandUseCase {
-KanjiToNumberConverter kanjiConverter
-CommandParser commandParser
+invoke(String text) Flow~CommandResult~
}
class KanjiToNumberConverter {
+convert(String text) String
}
class CommandParser {
-List~CommandDefinition~ commands
+parse(String text) CommandMatchResult
}
class CommandDefinition {
+CommandId id
+List~Regex~ patterns
+ICommandExecutor executor
}
class CommandId {
<<enum>>
GET_TIME_DATE
SET_ALARM_TIMER
CHANGE_SETTINGS
READ_MESSAGES
GET_SYSTEM_INFO
NAVIGATE_AI
UNKNOWN_FALLBACK
}
class ICommandExecutor {
<<interface>>
+execute(CommandMatchResult match) CommandResult
}
class NavigateAiExecutor {
-AiApiService apiService
-Context context
+execute(CommandMatchResult match) CommandResult
}
}
namespace Voice_Engine_Layer {
class IVoiceRecognizer {
<<interface>>
+StateFlow~VoiceRecognitionState~ state
+StateFlow~Float~ rmsDb
+startListening()
+stopListening()
+destroy()
}
class VoiceRecognitionManager {
-SpeechRecognizer speechRecognizer
}
class ITextToSpeech {
<<interface>>
+StateFlow~TtsStatus~ status
+speak(String text, Int queueMode) Unit
+stop()
+destroy()
}
class TextToSpeechManager {
-TextToSpeech ttsEngine
}
}
namespace Network_Layer {
class AiApiService {
<<interface>>
+navigate(AiApiRequest request) AiApiResponse
}
class AiApiRequest {
+String userId
+String queryText
+LocationData currentLocation
+Long timestamp
}
class AiApiResponse {
+String status
+String ttsMessage
+DestinationData destination
}
}
namespace Service_Layer {
class VoiceAssistantService {
-MediaSessionCompat mediaSession
+onStartCommand()
}
class AppNotificationListenerService {
+onNotificationPosted()
}
}
MainScreen ..> MainViewModel : Observe / Event
MainViewModel --> ExecuteCommandUseCase : Invoke
MainViewModel --> IVoiceRecognizer : Control
MainViewModel --> ITextToSpeech : Control
ExecuteCommandUseCase --> KanjiToNumberConverter : Pre-process
ExecuteCommandUseCase --> CommandParser : Use
CommandParser --> CommandDefinition : Contains
CommandDefinition --> ICommandExecutor : Delegates
ICommandExecutor <|.. NavigateAiExecutor : Implements
NavigateAiExecutor --> AiApiService : Call API
IVoiceRecognizer <|.. VoiceRecognitionManager : Implements
ITextToSpeech <|.. TextToSpeechManager : Implements
VoiceAssistantService --> IVoiceRecognizer : Trigger
```
#### 2.2.2 抽象インターフェース&データモデル定義 (Kotlin コード詳細)
##### 1. Voice Engine Layer インターフェース (`speak()` 戻り値 Unit 統一)
```kotlin
package com.example.voiceapp.voice
import kotlinx.coroutines.flow.StateFlow
sealed interface VoiceRecognitionState {
object Idle : VoiceRecognitionState
object Ready : VoiceRecognitionState
object Listening : VoiceRecognitionState
data class Success(val recognizedText: String) : VoiceRecognitionState
data class Error(val errorCode: Int, val message: String) : VoiceRecognitionState
}
interface IVoiceRecognizer {
val state: StateFlow<VoiceRecognitionState>
val rmsDb: StateFlow<Float>
fun startListening()
fun stopListening()
fun destroy()
}
sealed interface TtsStatus {
object Idle : TtsStatus
object Speaking : TtsStatus
object Completed : TtsStatus
data class Error(val message: String) : TtsStatus
}
interface ITextToSpeech {
val status: StateFlow<TtsStatus>
fun speak(text: String, queueMode: Int = 0): Unit
fun stop()
fun destroy()
}
```
##### 2. Network Layer API 定義 & データ構造
```kotlin
package com.example.voiceapp.data.api
import retrofit2.http.Body
import retrofit2.http.Header
import retrofit2.http.POST
data class LocationData(
val latitude: Double,
val longitude: Double
)
data class AiApiRequest(
val user_id: String,
val query_text: String,
val current_location: LocationData?,
val timestamp: Long
)
data class DestinationData(
val name: String,
val address: String,
val latitude: Double,
val longitude: Double
)
data class AiApiResponse(
val status: String,
val tts_message: String,
val destination: DestinationData?
)
data class AiApiErrorResponse(
val status: String,
val error_code: String,
val message: String
)
interface AiApiService {
@POST("api/v1/navigate")
suspend fun navigate(
@Header("Authorization") authHeader: String,
@Body request: AiApiRequest
): AiApiResponse
}
```
##### 3. Domain Layer インターフェース、漢数字コンバータ、コマンド構造
```kotlin
package com.example.voiceapp.domain.converter
class KanjiToNumberConverter {
private val kanjiMap = mapOf(
'零' to '0', '一' to '1', '二' to '2', '三' to '3', '四' to '4',
'五' to '5', '六' to '6', '七' to '7', '八' to '8', '九' to '9',
'0' to '0', '1' to '1', '2' to '2', '3' to '3', '4' to '4',
'5' to '5', '6' to '6', '7' to '7', '8' to '8', '9' to '9'
)
fun convert(input: String): String {
val sb = StringBuilder()
for (char in input) {
val replaced = kanjiMap[char]
if (replaced != null) {
sb.append(replaced)
} else {
sb.append(char)
}
}
return sb.toString()
}
}
```
```kotlin
package com.example.voiceapp.domain.model
enum class CommandId {
GET_TIME_DATE,
SET_ALARM_TIMER,
CHANGE_SETTINGS,
READ_MESSAGES,
GET_SYSTEM_INFO,
NAVIGATE_AI,
UNKNOWN_FALLBACK
}
data class CommandMatchResult(
val commandId: CommandId,
val matchedPattern: String,
val extractedParameters: Map<String, String>,
val rawText: String
)
data class CommandResult(
val commandId: CommandId,
val isSuccess: Boolean,
val responseText: String,
val data: Map<String, Any>? = null
)
interface ICommandExecutor {
suspend fun execute(matchResult: CommandMatchResult): CommandResult
}
data class CommandDefinition(
val id: CommandId,
val patterns: List<Regex>,
val executor: ICommandExecutor
)
```
```kotlin
package com.example.voiceapp.domain
import com.example.voiceapp.domain.converter.KanjiToNumberConverter
import com.example.voiceapp.domain.model.*
class CommandParser(private val definitions: List<CommandDefinition>) {
fun parse(text: String): CommandMatchResult {
for (def in definitions) {
for (pattern in def.patterns) {
val match = pattern.find(text)
if (match != null) {
val params = match.groups.mapIndexedNotNull { index, group ->
if (index > 0 && group != null) "param_$index" to group.value else null
}.toMap()
return CommandMatchResult(def.id, pattern.pattern, params, text)
}
}
}
return CommandMatchResult(CommandId.UNKNOWN_FALLBACK, "", emptyMap(), text)
}
}
class ExecuteCommandUseCase(
private val kanjiConverter: KanjiToNumberConverter,
private val commandParser: CommandParser
) {
suspend operator fun invoke(inputText: String): CommandResult {
val normalizedText = kanjiConverter.convert(inputText)
val matchResult = commandParser.parse(normalizedText)
return matchResult.executor.execute(matchResult)
}
}
```
##### 4. Presentation Layer 状態・メッセージモデル
```kotlin
package com.example.voiceapp.ui
import com.example.voiceapp.domain.model.CommandId
sealed interface VoiceAppState {
object Idle : VoiceAppState
data class Listening(val rmsDb: Float = 0f) : VoiceAppState
object Processing : VoiceAppState
data class Speaking(val text: String) : VoiceAppState
data class Error(val message: String, val errorCode: Int? = null) : VoiceAppState
}
enum class MessageSender { USER, SYSTEM }
data class ChatMessage(
val id: String = java.util.UUID.randomUUID().toString(),
val sender: MessageSender,
val text: String,
val timestamp: Long = System.currentTimeMillis(),
val commandId: CommandId? = null
)
```
#### 2.2.3 依存性注入 (DI: Hilt Framework Design)
```kotlin
package com.example.voiceapp.di
import android.content.Context
import com.example.voiceapp.data.api.AiApiService
import com.example.voiceapp.domain.*
import com.example.voiceapp.domain.converter.KanjiToNumberConverter
import com.example.voiceapp.domain.executor.*
import com.example.voiceapp.domain.model.*
import com.example.voiceapp.voice.*
import dagger.Module
import dagger.Provides
import dagger.hilt.InstallIn
import dagger.hilt.android.qualifiers.ApplicationContext
import dagger.hilt.components.SingletonComponent
import okhttp3.OkHttpClient
import retrofit2.Retrofit
import retrofit2.converter.gson.GsonConverterFactory
import java.util.concurrent.TimeUnit
import javax.inject.Singleton
@Module
@InstallIn(SingletonComponent::class)
object VoiceModule {
@Provides
@Singleton
fun provideVoiceRecognizer(
@ApplicationContext context: Context
): IVoiceRecognizer = VoiceRecognitionManager(context)
@Provides
@Singleton
fun provideTextToSpeech(
@ApplicationContext context: Context
): ITextToSpeech = TextToSpeechManager(context)
}
@Module
@InstallIn(SingletonComponent::class)
object NetworkModule {
@Provides
@Singleton
fun provideOkHttpClient(): OkHttpClient {
return OkHttpClient.Builder()
.connectTimeout(5000, TimeUnit.MILLISECONDS)
.readTimeout(10000, TimeUnit.MILLISECONDS)
.addInterceptor { chain ->
var request = chain.request()
val apiKey = "YOUR_API_KEY_HERE"
request = request.newBuilder()
.header("Authorization", "Bearer $apiKey")
.build()
// Exponential Backoff Retry (Max 3 Tries)
var response = chain.proceed(request)
var tryCount = 0
var backoffDelay = 1000L
while (!response.isSuccessful && tryCount < 3) {
tryCount++
Thread.sleep(backoffDelay)
backoffDelay *= 2
response.close()
response = chain.proceed(request)
}
response
}
.build()
}
@Provides
@Singleton
fun provideAiApiService(okHttpClient: OkHttpClient): AiApiService {
return Retrofit.Builder()
.baseUrl("https://api.example.com/")
.client(okHttpClient)
.addConverterFactory(GsonConverterFactory.create())
.build()
.create(AiApiService::class.java)
}
}
@Module
@InstallIn(SingletonComponent::class)
object DomainModule {
@Provides
@Singleton
fun provideKanjiToNumberConverter(): KanjiToNumberConverter {
return KanjiToNumberConverter()
}
@Provides
@Singleton
fun provideNavigateAiExecutor(
@ApplicationContext context: Context,
apiService: AiApiService
): NavigateAiExecutor {
return NavigateAiExecutor(context, apiService)
}
@Provides
@Singleton
fun provideCommandDefinitions(
@ApplicationContext context: Context,
navigateAiExecutor: NavigateAiExecutor
): List<CommandDefinition> {
return listOf(
CommandDefinition(
id = CommandId.NAVIGATE_AI,
patterns = listOf(Regex(".*(ナビ|行きたい|向かう|セット).*")),
executor = navigateAiExecutor
),
CommandDefinition(
id = CommandId.GET_TIME_DATE,
patterns = listOf(Regex(".*(今何時|何時ですか|日付|今日は何日).*")),
executor = GetTimeDateExecutor()
),
CommandDefinition(
id = CommandId.SET_ALARM_TIMER,
patterns = listOf(Regex(".*(\\d+)\\s*(分|秒|時間)タイマー.*|.*(\\d+)時にアラーム.*")),
executor = SetAlarmTimerExecutor(context)
),
CommandDefinition(
id = CommandId.CHANGE_SETTINGS,
patterns = listOf(Regex(".*(音量を|ダークモード|ライトモード).*")),
executor = ChangeSettingsExecutor(context)
),
CommandDefinition(
id = CommandId.READ_MESSAGES,
patterns = listOf(Regex(".*(メッセージ|未読|通知).*")),
executor = ReadMessagesExecutor(context)
),
CommandDefinition(
id = CommandId.GET_SYSTEM_INFO,
patterns = listOf(Regex(".*(バッテリー|電池|Wi-Fi|ネットワーク).*")),
executor = GetSystemInfoExecutor(context)
),
CommandDefinition(
id = CommandId.UNKNOWN_FALLBACK,
patterns = emptyList(),
executor = UnknownFallbackExecutor()
)
)
}
@Provides
@Singleton
fun provideCommandParser(definitions: List<CommandDefinition>): CommandParser {
return CommandParser(definitions)
}
@Provides
@Singleton
fun provideExecuteCommandUseCase(
kanjiConverter: KanjiToNumberConverter,
parser: CommandParser
): ExecuteCommandUseCase {
return ExecuteCommandUseCase(kanjiConverter, parser)
}
}
```
## 3. 回路・構造・システムの注意事項
### 3.1 物理・ハードウェア、電力、オーディオ、ネットワーク注意事項
1. **マイクフィードバック (AEC) & オーディオフォーカス制御**
- TTS 発話中のスピーカー音量をマイクが拾うループを完全防止するため、`ITextToSpeech` の `status` が `Speaking` の間は `IVoiceRecognizer` の `startListening()` 呼び出しをシステムレベルでロックする。
- `speak()` 呼出時に `AudioManager.requestAudioFocus(AUDIOFOCUS_GAIN_TRANSIENT_MAY_DUCK)` を要求し、他アプリの音量を自動ミュート/減衰させる。発話完了 (`onDone`) 時に `abandonAudioFocus()` を実行する。
2. **バッテリー消費抑制 & Bluetooth バックグラウンド起動制御 (BAL 対策)**
- Android 10+ (API 29+) のスリープ中・バックグラウンドからの Activity 起動制限 (Background Activity Launch restrictions) に対応するため、`VoiceAssistantService` を常駐 Foreground Service (`foregroundServiceType="microphone|location"`) として動作させる。
- Bluetooth ヘッドセットボタン (`Intent.ACTION_MEDIA_BUTTON`) 押下時は `MediaSessionCompat` の KeyEventListener でイベントを直接受信し、Service 内でマイク録音・音声認識を開始する。不要な Activity ダイレクト起動を行わないことで制限を回避し、かつ常時マイク監視を行わないため CPU の省電力サスペンドを維持する。
3. **通信信頼性・エラーハンドリング (Exponential Backoff & タイムアウト)**
- モバイル回線接続の不安定性に備え、REST API 通信には Connect Timeout 5,000ms, Read Timeout 10,000ms を設定。
- API 通信失敗時は Exponential Backoff (初期遅延 1,000ms, 倍率 2, 最大 3 回) にて自動リトライ。失敗継続時は `AiApiErrorResponse` (`LOCATION_NOT_FOUND` 等) を受託し、ユーザーへ親切なガイド音声を読み上げる。
### 3.2 システム制御シーケンス (シーケンス図)
#### 3.2.1 AI ナビゲーション実行 & 段階的フォールバックシーケンス
```mermaid
sequenceDiagram
autonumber
actor User as ユーザー
participant UI as Compose UI (MainScreen)
participant VM as MainViewModel
participant VR as IVoiceRecognizer
participant UC as ExecuteCommandUseCase
participant KC as KanjiToNumberConverter
participant CP as CommandParser
participant EX as NavigateAiExecutor
participant API as AiApiService (REST API)
participant TTS as ITextToSpeech
User->>UI: マイクボタンタップ / Bluetoothボタン
UI->>VM: onMicButtonClicked()
VM->>VR: startListening()
VR->>VM: state = Listening(rmsDb)
VM->>UI: StateFlow Update: VoiceAppState.Listening
User->>VR: 発話 (例:「仕事で行く三分後の〇〇会社に向かう」)
VR->>VM: state = Success("仕事で行く三分後の〇〇会社に向かう")
VM->>UI: StateFlow Update: VoiceAppState.Processing
VM->>VM: 会話履歴追加 (USER)
VM->>UC: invoke("仕事で行く三分後の〇〇会社に向かう")
UC->>KC: convert("仕事で行く三分後の〇〇会社に向かう")
KC-->>UC: 前処理結果 ("仕事で行く3分後の〇〇会社に向かう")
UC->>CP: parse("仕事で行く3分後の〇〇会社に向かう")
CP-->>UC: CommandMatchResult (NAVIGATE_AI)
UC->>EX: NavigateAiExecutor.execute()
EX->>API: POST /api/v1/navigate (Auth Header, 5s/10s Timeout)
API-->>EX: HTTP 200 (AiApiResponse: tts_message, destination)
EX->>EX: UTF-8 URL Encode: URLEncoder.encode(name, "UTF-8")
alt NaviCon インストール検出時
EX->>EX: NaviCon Intent (navicon://point?ll=lat,lng&title=encoded_name)
else NaviCon 未検出 且つ Google Maps インストール時
EX->>EX: Google Maps Intent (geo:lat,lng?q=encoded_name)
else いずれも未検出時
EX->>EX: Play Store Intent (market://details?id=jp.co.denso.navicon.user)
end
EX-->>UC: CommandResult (responseText = tts_message)
UC-->>VM: CommandResult 返却
VM->>TTS: speak(tts_message)
TTS->>VM: status = Speaking
VM->>UI: StateFlow Update: VoiceAppState.Speaking
TTS->>User: 音声出力
TTS->>VM: status = Completed
VM->>UI: StateFlow Update: VoiceAppState.Idle
```
#### 3.2.2 エラー発生時 3秒自動復帰シーケンス (delay 3000ms)
```mermaid
sequenceDiagram
autonumber
actor User as ユーザー
participant UI as MainScreen
participant VM as MainViewModel
participant VR as IVoiceRecognizer
participant TTS as ITextToSpeech
User->>VR: 発話不能 / ノイズ入力
VR->>VM: state = Error(ERROR_SPEECH_TIMEOUT)
VM->>UI: StateFlow Update: VoiceAppState.Error("音声が認識できませんでした")
VM->>TTS: speak("音声が認識できませんでした")
note over VM: MainViewModel 内で Coroutine delay(3000) 起動
VM->>VM: viewModelScope.launch { delay(3000); updateState(VoiceAppState.Idle) }
note over VM,UI: 3秒間エラーメッセージ・ダイアログを表示
VM->>UI: StateFlow Update: VoiceAppState.Idle
UI->>User: Idle 状態表示 (自動安全復帰完了)
```
#### 3.2.3 Bluetooth メディアボタン バックグラウンド起動シーケンス
```mermaid
sequenceDiagram
autonumber
actor User as ユーザー (ヘッドセット)
participant BT as Bluetooth Headset
participant SVC as VoiceAssistantService (Foreground)
participant MS as MediaSessionCompat.Callback
participant VR as IVoiceRecognizer
participant VM as MainViewModel
User->>BT: メディアボタン押下
BT->>MS: Intent.ACTION_MEDIA_BUTTON
MS->>SVC: onMediaButtonEvent()
SVC->>SVC: チェック: 常駐 Foreground Service 動作中
SVC->>VR: startListening()
VR->>VM: state = Listening
VM->>User: 音声認識開始 (画面オフ状態でも制御継続)
```
### 3.3 ナビゲーション連携 (NaviCon & Google Maps) & URL エンコード仕様
1. **NaviCon 連携 (メイン軸 / 完全無料)**
- API 応答から受信した目的地名 (`name`) に必ず `URLEncoder.encode(name, "UTF-8")` 処理を行う。
- スキームフォーマット: `navicon://point?ll={latitude},{longitude}&title={encoded_name}`
2. **Google Maps 連携 (フォールバック 1)**
- スキームフォーマット: `geo:{latitude},{longitude}?q={encoded_name}`
3. **Google Play ストア誘導 (フォールバック 2)**
- スキームフォーマット: `market://details?id=jp.co.denso.navicon.user`
4. **パッケージ確認ロジック (フォールバック判定順序)**
- `PackageManager.getPackageInfo("jp.co.denso.navicon.user", 0)` を試行。
- 存在する場合 -> NaviCon Intent 発行。
- 存在しない場合 -> 「NaviConが未検出のため、Google Mapsで表示します」と TTS 発話後、Google Maps Intent 発行。
- Google Maps も未検出の場合 -> Playストア Intent 発行。
### 3.4 ディレクトリ構造とソースファイル配置設計
ルートディレクトリ構成からアプリ内部のパッケージ構成に至る完全なディレクトリ構造は以下の通り。
```
.
├── .gitea/
│ └── workflows/
│ └── build.yaml # Gitea Runner CI/CD ワークフロー
├── docker/
│ ├── Dockerfile # Android ビルド用 Dockerfile (Ubuntu 22.04 + JDK 17 + Android SDK)
│ └── entrypoint.sh # コンテナ起動初期化スクリプト
├── docker-compose.yml # ローカル Docker ビルド定義
├── app/
│ ├── build.gradle.kts # アプリ依存関係設定
│ └── src/
│ ├── main/
│ │ ├── java/com/example/voiceapp/
│ │ │ ├── VoiceApplication.kt # Application クラス (@HiltAndroidApp)
│ │ │ ├── MainActivity.kt # シングルアクティビティ
│ │ │ ├── data/
│ │ │ │ └── api/
│ │ │ │ ├── AiApiService.kt
│ │ │ │ └── models/ # AiApiRequest, AiApiResponse, AiApiErrorResponse
│ │ │ ├── di/
│ │ │ │ ├── VoiceModule.kt
│ │ │ │ ├── NetworkModule.kt
│ │ │ │ └── DomainModule.kt
│ │ │ ├── domain/
│ │ │ │ ├── converter/
│ │ │ │ │ └── KanjiToNumberConverter.kt
│ │ │ │ ├── executor/
│ │ │ │ │ ├── ICommandExecutor.kt
│ │ │ │ │ ├── GetTimeDateExecutor.kt
│ │ │ │ │ ├── SetAlarmTimerExecutor.kt
│ │ │ │ │ ├── ChangeSettingsExecutor.kt
│ │ │ │ │ ├── ReadMessagesExecutor.kt
│ │ │ │ │ ├── GetSystemInfoExecutor.kt
│ │ │ │ │ ├── NavigateAiExecutor.kt
│ │ │ │ │ └── UnknownFallbackExecutor.kt
│ │ │ │ ├── model/
│ │ │ │ │ ├── CommandId.kt
│ │ │ │ │ ├── CommandMatchResult.kt
│ │ │ │ │ └── CommandResult.kt
│ │ │ │ ├── CommandDefinition.kt
│ │ │ │ ├── CommandParser.kt
│ │ │ │ └── ExecuteCommandUseCase.kt
│ │ │ ├── voice/
│ │ │ │ ├── IVoiceRecognizer.kt
│ │ │ │ ├── VoiceRecognitionManager.kt
│ │ │ │ ├── ITextToSpeech.kt
│ │ │ │ └── TextToSpeechManager.kt
│ │ │ ├── service/
│ │ │ │ ├── VoiceAssistantService.kt
│ │ │ │ └── AppNotificationListenerService.kt
│ │ │ └── ui/
│ │ │ ├── VoiceAppState.kt
│ │ │ ├── ChatMessage.kt
│ │ │ ├── MainViewModel.kt
│ │ │ └── components/
│ │ │ ├── MainScreen.kt
│ │ │ ├── StatusBanner.kt
│ │ │ ├── WaveformIndicator.kt
│ │ │ ├── ConversationHistory.kt
│ │ │ ├── ControlBar.kt
│ │ │ └── PermissionDialog.kt
│ │ └── AndroidManifest.xml
│ └── test/
│ └── java/com/example/voiceapp/
│ ├── domain/
│ │ ├── KanjiToNumberConverterTest.kt
│ │ ├── CommandParserTest.kt
│ │ └── ExecuteCommandUseCaseTest.kt
│ └── ui/
│ └── MainViewModelTest.kt
└── README.md
```
### 3.5 AndroidManifest.xml 権限設定・サービス設定設計
Android OS 8.0+ (API Level 26) から Android 14 (API Level 34) までの互換性および必須パーミッション・フォアグラウンドサービス宣言を網羅する `AndroidManifest.xml` の完全記述。
```xml
<?xml version="1.0" encoding="utf-8"?>
<manifest xmlns:android="http://schemas.android.com/apk/res/android"
package="com.example.voiceapp">
<!-- 必須パーミッション宣言 -->
<!-- 1. 音声入力権限 (危険権限) -->
<uses-permission android:name="android.permission.RECORD_AUDIO" />
<!-- 2. AI サーバー通信・ネットワーク確認権限 -->
<uses-permission android:name="android.permission.INTERNET" />
<uses-permission android:name="android.permission.ACCESS_NETWORK_STATE" />
<!-- 3. 位置情報権限 (AI ナビゲーション現在地付与用) -->
<uses-permission android:name="android.permission.ACCESS_FINE_LOCATION" />
<uses-permission android:name="android.permission.ACCESS_COARSE_LOCATION" />
<!-- 4. オーディオ設定(音量変更・フォーカス要求)権限 -->
<uses-permission android:name="android.permission.MODIFY_AUDIO_SETTINGS" />
<!-- 5. アラーム・タイマーインテント設定権限 (Android 12+ API 31+) -->
<uses-permission android:name="android.permission.SCHEDULE_EXACT_ALARM" />
<!-- 6. 通知出力権限 (Android 13+ API 33+) -->
<uses-permission android:name="android.permission.POST_NOTIFICATIONS" />
<!-- 7. 未読メッセージ通知受託権限 -->
<uses-permission android:name="android.permission.BIND_NOTIFICATION_LISTENER_SERVICE" />
<!-- 8. Android 14 (API 34) 必須フォアグラウンドサービス権限 -->
<uses-permission android:name="android.permission.FOREGROUND_SERVICE" />
<uses-permission android:name="android.permission.FOREGROUND_SERVICE_MICROPHONE" />
<uses-permission android:name="android.permission.FOREGROUND_SERVICE_LOCATION" />
<!-- ハードウェア機能制限の設定 (マイク必須設定) -->
<uses-feature
android:name="android.hardware.microphone"
android:required="true" />
<uses-feature
android:name="android.hardware.location.gps"
android:required="false" />
<application
android:name=".VoiceApplication"
android:allowBackup="true"
android:icon="@mipmap/ic_launcher"
android:label="音声操作ナビ端末"
android:roundIcon="@mipmap/ic_launcher_round"
android:supportsRtl="true"
android:theme="@style/Theme.VoiceControlledDevice">
<!-- シングルアクティビティ構成 -->
<activity
android:name=".MainActivity"
android:exported="true"
android:launchMode="singleTop"
android:screenOrientation="portrait"
android:windowSoftInputMode="adjustResize">
<intent-filter>
<action android:name="android.intent.action.MAIN" />
<category android:name="android.intent.category.LAUNCHER" />
</intent-filter>
<!-- 音声アシスタント呼び出しインテント -->
<intent-filter>
<action android:name="android.intent.action.VOICE_COMMAND" />
<category android:name="android.intent.category.DEFAULT" />
</intent-filter>
</activity>
<!-- 常駐フォアグラウンドサービス (Android 14 API 34 対応) -->
<service
android:name=".service.VoiceAssistantService"
android:exported="false"
android:foregroundServiceType="microphone|location" />
<!-- 未読メッセージ読み上げ通知受託サービス -->
<service
android:name=".service.AppNotificationListenerService"
android:label="メッセージ読み上げサービス"
android:permission="android.permission.BIND_NOTIFICATION_LISTENER_SERVICE"
android:exported="true">
<intent-filter>
<action android:name="android.service.notification.NotificationListenerService" />
</intent-filter>
</service>
<!-- SpeechRecognizer, NaviCon, Google Maps パッケージクエリ設定 (Android 11+ API 30+) -->
<queries>
<intent>
<action android:name="android.speech.RecognitionService" />
</intent>
<intent>
<action android:name="android.intent.action.TTS_SERVICE" />
</intent>
<package android:name="jp.co.denso.navicon.user" />
<package android:name="com.google.android.apps.maps" />
</queries>
</application>
</manifest>
```
### 3.6 Gitea CI/CD Docker Runner 環境・ビルドパイプライン設計
1. **Docker Runner ビルドモデル**
- CI/CD パイプラインは Gitea Act Runner の Docker Runner モードにて動的コンテナ環境を生成して実行する。
- ホストOSに依存せず、コンテナイメージ (`docker/Dockerfile`) 内で JDK 17, Android SDK (API 34 Build-tools), Gradle 8.x 環境を完全にカプセル化する。
2. **Gitea ワークフロー定義 (`.gitea/workflows/build.yaml`)**
```yaml
name: Android CI Build & Test
on:
push:
branches: [ main, develop ]
pull_request:
branches: [ main, develop ]
jobs:
build-and-test:
runs-on: ubuntu-latest
container:
image: voice-app-build:latest
steps:
- name: Checkout Repository
uses: actions/checkout@v3
- name: Run Unit Tests
run: ./gradlew testDebugUnitTest --stacktrace
- name: Assemble Debug APK
run: ./gradlew assembleDebug --stacktrace
- name: Upload Artifacts
uses: actions/upload-artifact@v3
with:
name: debug-apk
path: app/build/outputs/apk/debug/app-debug.apk
```
3. **ローカル検証モデル (`docker-compose.yml`)**
- ローカル環境でも `docker-compose run --rm build-env ./gradlew testDebugUnitTest` で CI と同一のテスト・ビルド環境を実行可能とする。