Patrick Baumgartner 3c80f37443 Add voice recognition optional feature and update .gitignore vor 1 Jahr
..
README.md d9faf40c75 In between commit vor 1 Jahr
__init__.py d9faf40c75 In between commit vor 1 Jahr
audio_buffer.py d9faf40c75 In between commit vor 1 Jahr
audio_features.py d9faf40c75 In between commit vor 1 Jahr
config.py d9faf40c75 In between commit vor 1 Jahr
detection_engine.py d9faf40c75 In between commit vor 1 Jahr
detector.py 3c80f37443 Add voice recognition optional feature and update .gitignore vor 1 Jahr
examples.py d9faf40c75 In between commit vor 1 Jahr
model_loader.py 3c80f37443 Add voice recognition optional feature and update .gitignore vor 1 Jahr
optimization.py d9faf40c75 In between commit vor 1 Jahr

README.md

Trixy Wakeword Detection System

A comprehensive, production-ready wakeword detection system for the Trixy voice assistant. This system provides real-time detection of custom wakewords using PyTorch models with advanced features for low-latency, high-accuracy performance.

Features

Core Capabilities

  • Dual Wakeword Detection: Supports both "custom" and "system_command" wakewords
  • Real-time Processing: Optimized for low-latency streaming audio processing
  • Password-Protected Models: Secure model loading with ZIP-based encryption
  • Event System Integration: Seamless integration with Trixy's event-driven architecture
  • Performance Monitoring: Comprehensive metrics and optimization tools

Audio Processing

  • Log-Mel Spectrograms: High-quality feature extraction optimized for speech
  • Configurable Parameters: 20-40 mel filterbanks, 10ms shift, 16kHz audio
  • Real-time Buffering: Circular buffers with configurable overlap
  • Noise Handling: Built-in audio normalization and preprocessing

Model Support

  • RepCNN Architecture: Efficient reparameterizable CNN models
  • Multiple Variants: Standard, Improved, and Lightweight model options
  • Model Optimization: TorchScript, quantization, and compilation support
  • Device Flexibility: CPU, CUDA, and MPS device support

Performance Optimization

  • Temporal Filtering: Reduces false positives with temporal consistency
  • Confidence Calibration: Improved probability estimates
  • Memory Management: Efficient memory usage with garbage collection
  • Batch Processing: Optional batch inference for throughput optimization

Quick Start

Basic Usage

from trixy_core.ml.wakeword import WakewordDetector, create_default_config

# Create configuration
config = create_default_config()
config.model_config.model_path = "path/to/your/model.pth"
config.model_config.model_password = "your_password"

# Initialize detector with event handler
detector = WakewordDetector(config, event_handler)

# Start detection
detector.start()

# Process audio data (numpy array, 16kHz, mono)
results = detector.process_audio(audio_data)

# Stop when done
detector.stop()

Factory Function

from trixy_core.ml.wakeword import create_wakeword_detector

# Simplified creation
detector = create_wakeword_detector(
    model_path="path/to/model.pth",
    model_password="password",
    event_handler=event_handler,
    custom_threshold=0.8,
    system_command_threshold=0.9
)

Context Manager

with WakewordDetector(config, event_handler) as detector:
    # Process audio
    results = detector.process_audio(audio_data)
    # Automatically stopped when exiting context

Configuration

Performance Profiles

from trixy_core.ml.wakeword import (
    create_lightweight_config,
    create_high_performance_config,
    auto_tune_config
)

# For resource-constrained devices
lightweight_config = create_lightweight_config()

# For high-performance systems
high_perf_config = create_high_performance_config()

# Auto-tuned for specific requirements
tuned_config = auto_tune_config(
    base_config,
    target_latency_ms=50.0,
    available_memory_mb=256.0
)

Custom Configuration

from trixy_core.ml.wakeword import WakewordConfig, ModelConfig, AudioConfig

config = WakewordConfig(
    model_config=ModelConfig(
        model_path="models/wakeword/custom.pth",
        model_password="secure_password",
        device="cuda",
        use_half_precision=True
    ),
    audio_config=AudioConfig(
        sample_rate=16000,
        chunk_duration=1.0,
        overlap_ratio=0.5
    ),
    detection_config=DetectionConfig(
        custom_threshold=0.75,
        system_command_threshold=0.85,
        use_temporal_smoothing=True
    )
)

Event Integration

The system integrates with Trixy's event system to trigger wakeword_received events:

from trixy_core.events.decorators import TrixyEvent

class MyPlugin(TrixyPlugin):
    @TrixyEvent(["wakeword_received"])
    def on_wakeword_detected(self, event_name, event_data):
        wakeword_type = event_data.wakeword_type  # "custom" or "system_command"
        confidence = event_data.confidence
        satellite_id = event_data.satellite_info.satellite_id
        
        if wakeword_type == "system_command":
            # Handle admin commands
            self.handle_system_command()
        else:
            # Handle regular wakeword
            self.start_conversation()

Model Management

Loading Models

from trixy_core.ml.wakeword import SecureModelLoader

loader = SecureModelLoader(device=torch.device("cuda"))
model, metadata = loader.load_model(
    model_path="model.pth",
    password="password",
    verify_hash=True
)

print(f"Loaded model: {metadata.model_name}")
print(f"Classes: {metadata.class_labels}")
print(f"Input shape: {metadata.input_shape}")

Model Metadata

Models include comprehensive metadata:

metadata = ModelMetadata(
    model_name="Custom Wakeword v2.1",
    model_version="2.1.0",
    model_type="wakeword",
    architecture="ImprovedRepCNN",
    num_classes=3,
    class_labels=["custom", "system_command", "negative"],
    final_accuracy=0.956,
    sample_rate=16000,
    n_mels=40,
    time_frames=151
)

Performance Monitoring

Real-time Monitoring

from trixy_core.ml.wakeword import PerformanceMonitor

monitor = PerformanceMonitor(
    monitoring_interval=1.0,
    memory_warning_threshold=512.0,
    latency_warning_threshold=100.0
)

monitor.start_monitoring()
# ... process audio ...
monitor.stop_monitoring()

stats = monitor.get_performance_summary()
print(f"Average latency: {stats['avg_inference_latency_ms']:.1f}ms")

Model Optimization

from trixy_core.ml.wakeword import ModelOptimizer, optimize_for_deployment

# Benchmark model performance
optimizer = ModelOptimizer()
benchmark_results = optimizer.benchmark_model(
    model, input_shape=(1, 40, 151), device=device
)

# Optimize for deployment
optimized_model, report = optimize_for_deployment(
    model, 
    config={'enable_torchscript': True, 'enable_compilation': True},
    example_input
)

Testing and Validation

Comprehensive Testing

from trixy_core.ml.wakeword import WakewordTester, AudioSimulator

# Create test environment
audio_sim = AudioSimulator(sample_rate=16000)
tester = WakewordTester(detector, audio_sim)

# Run various tests
basic_result = tester.run_basic_test()
perf_result = tester.run_performance_test(duration_seconds=30.0)
accuracy_result = tester.run_accuracy_test(test_cases)
latency_result = tester.run_latency_test(num_iterations=100)

# Generate report
report = tester.get_test_report()
tester.save_test_report("test_results.json")

Audio Simulation

# Generate test audio
silence = audio_sim.generate_silence(1.0)
noise = audio_sim.generate_white_noise(1.0, amplitude=0.1)
synthetic_wakeword = audio_sim.generate_synthetic_wakeword(1.0)
test_sequence = audio_sim.create_test_sequence()

Architecture Overview

System Components

┌─────────────────┐    ┌──────────────────┐    ┌─────────────────┐
│   Audio Input   │───▶│   Audio Buffer   │───▶│ Feature Extract │
└─────────────────┘    └──────────────────┘    └─────────────────┘
                                                          │
┌─────────────────┐    ┌──────────────────┐    ┌─────────────────┐
│  Event System   │◀───│ Detection Engine │◀───│   Model Loader  │
└─────────────────┘    └──────────────────┘    └─────────────────┘

Processing Pipeline

  1. Audio Input: Raw 16kHz mono audio streams
  2. Audio Buffering: Circular buffer with configurable overlap
  3. Feature Extraction: Log-Mel spectrogram generation
  4. Model Inference: RepCNN-based wakeword classification
  5. Post-Processing: Confidence calibration and temporal filtering
  6. Event Triggering: Integration with Trixy event system

Model Architecture

The system uses RepCNN (Reparameterizable CNN) models optimized for wakeword detection:

  • Input: Log-Mel spectrograms (1×40×151 for 1s audio)
  • Architecture: Efficient CNN with reparameterizable blocks
  • Output: 3-class classification (custom, system_command, negative)
  • Optimization: TorchScript compilation, quantization support

Configuration Files

Default Configuration

{
  "model_config": {
    "model_path": "models/wakeword/default/model.pth",
    "model_password": "",
    "device": "auto",
    "use_model_optimization": true
  },
  "audio_config": {
    "sample_rate": 16000,
    "chunk_duration": 1.0,
    "overlap_ratio": 0.5,
    "spectrogram_config": {
      "n_mels": 40,
      "time_frames": 151,
      "hop_length": 160
    }
  },
  "detection_config": {
    "custom_threshold": 0.7,
    "system_command_threshold": 0.8,
    "use_temporal_smoothing": true
  }
}

Environment Variables

export TRIXY_WAKEWORD_MODEL_PATH="/path/to/model.pth"
export TRIXY_WAKEWORD_MODEL_PASSWORD="secure_password"
export TRIXY_WAKEWORD_DEVICE="cuda"
export TRIXY_WAKEWORD_CUSTOM_THRESHOLD="0.75"
export TRIXY_WAKEWORD_SYSTEM_THRESHOLD="0.85"
export TRIXY_WAKEWORD_DEBUG_MODE="true"

Performance Characteristics

Typical Performance

  • Inference Latency: 10-50ms on modern hardware
  • Memory Usage: 50-200MB depending on configuration
  • CPU Usage: 5-15% on single core
  • Accuracy: >95% on clean audio, >90% with noise

Optimization Guidelines

  1. Low Latency: Use lightweight config, disable temporal smoothing
  2. High Accuracy: Enable temporal smoothing, use improved model
  3. Low Memory: Reduce buffer size, enable memory management
  4. High Throughput: Enable batch processing, use GPU acceleration

Integration Examples

Satellite Client Integration

class TrixyClient:
    def __init__(self):
        config = create_default_config()
        config.satellite_id = self.get_satellite_id()
        self.wakeword_detector = WakewordDetector(config, self.event_handler)
    
    def start_audio_processing(self):
        self.wakeword_detector.start()
        # Start audio capture thread
        self.audio_thread = threading.Thread(target=self.audio_capture_loop)
        self.audio_thread.start()
    
    def audio_capture_loop(self):
        while self.running:
            audio_data = self.capture_audio_chunk()
            self.wakeword_detector.process_audio(audio_data)

Plugin Integration

class WakewordPlugin(TrixyPlugin):
    def __init__(self):
        super().__init__()
        self.detector = create_wakeword_detector(
            config_path=self.config['wakeword_config'],
            event_handler=self.application.event_handler
        )
    
    def on_plugin_start(self):
        self.detector.start()
    
    def on_plugin_stop(self):
        self.detector.stop()
    
    @TrixyEvent(["audio_chunk_received"])
    def process_audio_chunk(self, event_name, event_data):
        self.detector.process_audio(event_data.audio_data)

Troubleshooting

Common Issues

  1. Model Loading Errors

    • Check file path and password
    • Verify model format and metadata
    • Ensure sufficient permissions
  2. High Latency

    • Reduce chunk duration
    • Disable temporal smoothing
    • Use lighter model variant
  3. High Memory Usage

    • Reduce buffer duration
    • Enable memory management
    • Use half precision
  4. Low Accuracy

    • Adjust confidence thresholds
    • Enable temporal smoothing
    • Check audio quality and format

Debug Mode

Enable debug mode for detailed logging:

config.debug_mode = True
config.save_debug_audio = True
config.debug_output_path = "/tmp/wakeword_debug"

Contributing

When contributing to the wakeword detection system:

  1. Follow the existing code structure and patterns
  2. Add comprehensive tests for new features
  3. Update documentation and examples
  4. Ensure performance regression testing
  5. Validate with multiple model architectures

License

This wakeword detection system is part of the Trixy voice assistant project.