Setting the file. One moment. CHECKLIST · Agentic UX Design Relationship Centric Interfaces · bencium/bencium-marketplace · Skills Docs
This document provides practical checklists, worksheets, and audit tools for implementing relationship-centric design.
Relationship UX Audit Checklist
Use this checklist to evaluate existing interfaces or plan new ones.
Current State Assessment:
- Not done: System remembers user preferences (basic: theme, language)
- Not done: System tracks user behavior patterns
- Not done: System maintains context across sessions
- Not done: System recognizes user’s emotional state (frustration, satisfaction)
- Not done: System understands temporal patterns (time of day, day of week)
- Not done: System learns from user’s actual behavior (not just stated preferences)
- Not done: System maintains awareness of user’s ongoing goals
- Not done: System provides cross-device context continuity
- Not done: What user context is currently lost between sessions?
- Not done: What behavioral patterns would be valuable to track?
- Not done: What temporal patterns affect user behavior?
- Not done: What emotional states impact user experience?
- Not done: High: Essential context that users complain about losing
- Not done: Medium: Patterns that would improve experience noticeably
- Not done: Low: Nice-to-have personalization
Current State Assessment:
- Not done: System explains its reasoning for suggestions/decisions
- Not done: System shows confidence levels
- Not done: System allows users to adjust autonomy levels
- Not done: System has different trust stages (transparency → selective → autonomous)
- Not done: System provides undo/correction mechanisms
- Not done: System learns from user corrections
- Not done: System has trust recovery protocols for mistakes
- Not done: System escalates uncertain decisions appropriately
- Not done: How transparent is system reasoning currently?
- Not done: Can users control autonomy levels?
- Not done: What happens when system makes mistakes?
- Not done: How does trust evolve over time (or does it)?
- Not done: Define what should always be transparent (high-stakes, uncertain)
- Not done: Define what can become autonomous (routine, high-confidence)
- Not done: Design transition criteria between stages
- Not done: Create user controls for trust progression
Current Metrics Assessment:
- Not done: We measure: session duration, page views, conversion rates (traditional)
- Not done: We measure: relationship quality indicators
- Not done: We measure: compounding value over time
- Not done: We measure: context accuracy
- Not done: We measure: democratic alignment / ethical boundaries
- Not done: We track metrics longitudinally (weeks/months, not just sessions)
- Not done: We compare Month 1 vs. Month 6 experience quality
- Not done: What traditional metrics are misleading for our use case?
- Not done: What relationship metrics would better indicate success?
- Not done: How do we measure improvement over time?
- Not done: What longitudinal tracking do we need?
New Metrics to Implement:
- Not done: Relationship Quality: Trust scores, delegation comfort
- Not done: Compounding Value: Time-to-success improvement, capability expansion
- Not done: Context Accuracy: Intent prediction, preference matching
- Not done: Democratic Alignment: Value alignment, boundary respect
Current State Assessment:
- Not done: System understands user’s ongoing goals
- Not done: System provides proactive suggestions (not just reactive)
- Not done: System and user co-create plans together
- Not done: System adapts interface based on usage patterns
- Not done: System learns from user’s path choices
- Not done: System generates alternative paths dynamically
- Not done: System recognizes when user is stuck/frustrated
- Not done: System offers help at appropriate moments (not intrusive)
- Not done: How well does system understand user goals?
- Not done: Is system proactive or only reactive?
- Not done: Do users and system collaborate on planning?
- Not done: Does interface adapt to individual users?
Collaborative Features to Add:
- Not done: Goal capture and tracking interface
- Not done: Proactive suggestion engine
- Not done: Co-creation workspace (human + AI contributions)
- Not done: Adaptive UI elements
- Not done: Learning feedback mechanisms
Current State Assessment:
- Not done: Users can see what system remembers about them
- Not done: Users can control what gets remembered
- Not done: Users can forget/delete specific memories
- Not done: Users can export their data
- Not done: System has clear retention policies
- Not done: System scrubs PII appropriately
- Not done: System respects user boundaries explicitly
- Not done: System explains data usage transparently
- Not done: What memory controls do users currently have?
- Not done: Can users see and manage what’s remembered?
- Not done: Are privacy policies clear and actionable?
- Not done: Do users understand data retention?
Privacy Controls to Implement:
- Not done: Memory visualization interface
- Not done: Granular forgetting controls
- Not done: Data export functionality
- Not done: Clear retention policy UI
- Not done: Boundary setting interface
A structured 3-day sprint to design your memory architecture.
Morning: Behavioral Inventory (3 hours)
-
List all user interactions (30 min)
- What actions can users take?
- What choices do users make?
- What paths do users follow?
-
Identify valuable patterns (60 min)
- Which patterns indicate user goals?
- Which patterns indicate frustration?
- Which patterns indicate satisfaction?
- Which patterns evolve over time?
-
Map temporal dimensions (30 min)
- What changes by time of day?
- What changes by day of week?
- What changes by season/context?
-
Define context signals (60 min)
- What indicates user’s current state?
- What indicates user’s goals?
- What indicates user’s constraints?
Afternoon: Privacy & Retention Design (3 hours)
-
Categorize memory types (45 min)
- Behavioral patterns
- Explicit preferences
- Ongoing goals
- Historical trends
- Sensitive information
-
Define retention policies (45 min)
- What must be kept?
- What should expire?
- What should users control?
- What requires special handling?
-
Design user controls (90 min)
- Memory visualization interface
- Forgetting controls
- Export functionality
- Boundary settings
Day 1 Deliverable: Memory inventory and privacy framework
Morning: Data Structures (3 hours)
-
Design event schema (60 min)
-
Design pattern schema (60 min)
-
Design memory graph (60 min)
Afternoon: Implementation Planning (3 hours)
-
Storage strategy (45 min)
- Hot memory (recent, always loaded)
- Warm memory (patterns, load on demand)
- Cold memory (historical, analysis only)
-
Pattern detection algorithms (90 min)
- Frustration indicators
- Success patterns
- Temporal patterns
- Preference evolution
-
Privacy implementation (45 min)
- PII scrubbing
- Encryption
- Access controls
- Audit logging
Day 2 Deliverable: Complete memory architecture specification
Morning: Metrics Design (3 hours)
-
Select relationship metrics (45 min)
Choose 2-3 from each category:
- Relationship Quality
- Compounding Value
- Context Accuracy
- Democratic Alignment
-
Define measurement methods (90 min)
For each metric:
- How to calculate?
- What data needed?
- What indicates success?
- What indicates problems?
-
Design metric visualization (45 min)
- Dashboard layout
- Trend displays
- Alert thresholds
- User-facing vs. internal metrics
Afternoon: Testing Strategy (3 hours)
-
Longitudinal test plan (60 min)
- Week 1 tests (onboarding, transparency)
- Month 1 tests (pattern learning, trust evolution)
- Month 3 tests (compounding value, autonomy)
- Month 6 tests (relationship maturity)
-
Success criteria (60 min)
- Week 1: Can system explain reasoning? Do users understand?
- Month 1: Are patterns being detected? Is trust evolving?
- Month 3: Is experience improving? Are metrics trending positive?
- Month 6: Is compounding value evident? Are users delegating comfortably?
-
Risk mitigation (60 min)
- Privacy risks and mitigations
- Trust violation risks and recovery protocols
- Metric degradation and intervention triggers
Day 3 Deliverable: Metrics specification and longitudinal test plan
Use this worksheet to design trust evolution for your specific domain.
High-stakes decisions (ALWAYS transparent):
- Decision: _______________________________________________
- Why high-stakes: ________________________________________
- What to explain: _________________________________________
Uncertain decisions (ALWAYS show confidence):
- Decision: _______________________________________________
- Uncertainty source: ______________________________________
- Confidence threshold: ____________________________________
Routine decisions (CAN become autonomous):
- Decision: _______________________________________________
- Why routine: ____________________________________________
- Autonomy criteria: _______________________________________
Stage 1: Transparency Phase (Weeks 1-X)
- Not done: Understand system reasoning
- Not done: See confidence levels
- Not done: Learn system capabilities
- Not done: Build initial trust
- Not done: Explain all suggestions
- Not done: Show data sources
- Not done: Display alternatives considered
- Not done: Highlight uncertainties
- Not done: User accepts suggestions >60%
- Not done: User requests explanations <40%
- Not done: User indicates comfort with system
Stage 2: Selective Disclosure Phase (Weeks X-Y)
- Not done: See reasoning for important decisions
- Not done: Trust system for routine actions
- Not done: Understand when system is uncertain
- Not done: Full transparency for high-stakes decisions
- Not done: Confidence indicators for medium-stakes
- Not done: Quiet execution for routine tasks
- Not done: Clear escalation for uncertainties
- Not done: User accepts suggestions >75%
- Not done: User comfortable with routine autonomy
- Not done: User delegates specific categories
Stage 3: Autonomous Action Phase (Month Y+)
- Not done: Trust system to act independently
- Not done: Easy correction mechanisms
- Not done: Transparency on demand
- Not done: Control over autonomy levels
- Not done: Act autonomously for delegated categories
- Not done: Subtle notifications for actions taken
- Not done: Escalate uncertainties appropriately
- Not done: Learn from corrections
- Not done: Consistent with user values
- Not done: Clear undo mechanisms
- Not done: Explain on demand
- Not done: Adjust autonomy based on feedback
- Confidence meter: How to display? ____________________________
- Reasoning toggle: Where to place? ____________________________
- Autonomy controls: How to adjust? ____________________________
- Trust score: Show to user? ____________________________
Trust evolution feedback:
- How does user know trust is evolving? ________________________
- How does user see autonomy increasing? ______________________
- How does user control progression? __________________________
When system makes mistake:
-
Acknowledge (within X hours/days): ______________________
Template: “I made a suboptimal decision about [X] because [Y]”
-
Explain what went wrong: ________________________________
- Wrong assumption: _______________________________________
- What system learned: ____________________________________
-
Offer corrections:
- Option A: _______________________________________________
- Option B: _______________________________________________
- User’s choice: __________________________________________
-
Adjust trust level:
- Reduce autonomy in category: ____________________________
- For duration: ___________________________________________
- Re-evaluation criteria: _________________________________
-
Update model:
- Pattern learned: ________________________________________
- New guardrails: _________________________________________
Relationship Quality Metrics:
- Not done: Calculation method: _____________________________________
- Not done: Data sources: ___________________________________________
- Not done: Success threshold: ______________________________________
Option 2: Delegation Comfort
- Not done: Measurement: ____________________________________________
- Not done: Categories to track: ____________________________________
- Not done: Target: _________________________________________________
- Not done: What counts as override: ________________________________
- Not done: Acceptable rate: ________________________________________
- Not done: Alert threshold: ________________________________________
Compounding Value Metrics:
Option 1: Time-to-Success Improvement
- Not done: Baseline measurement: ___________________________________
- Not done: Current measurement: ____________________________________
- Not done: Improvement rate: _______________________________________
Option 2: Capability Expansion
- Not done: Features used at baseline: _____________________________
- Not done: Features used currently: ________________________________
- Not done: Adoption rate: __________________________________________
Option 3: Outcome Quality
- Not done: Quality measurement: ____________________________________
- Not done: Baseline vs. current: ___________________________________
- Not done: Improvement trend: ______________________________________
Context Accuracy Metrics:
Option 1: Intent Prediction Accuracy
- Not done: How to measure intent: __________________________________
- Not done: Prediction vs. actual: __________________________________
- Not done: Target accuracy: ________________________________________
Option 2: Preference Matching
- Not done: Recommendation acceptance rate: _________________________
- Not done: Top choice hit rate: ____________________________________
- Not done: Target: _________________________________________________
Option 3: Timing Accuracy
- Not done: Suggestion timing evaluation: __________________________
- Not done: User feedback on timing: ________________________________
- Not done: Target: _________________________________________________
Democratic Alignment Metrics:
Option 1: Value Alignment Score
- Not done: Constitution/values defined: ___________________________
- Not done: Violation detection: ____________________________________
- Not done: Target: _________________________________________________
Option 2: Boundary Respect
- Not done: Boundaries defined: _____________________________________
- Not done: Violation tracking: _____________________________________
- Not done: Target: _________________________________________________
Option 3: Fairness Metric
- Not done: Segments to compare: ____________________________________
- Not done: Fairness calculation: ___________________________________
- Not done: Target: _________________________________________________
- Metrics to capture: _________________________________________
- Measurement method: __________________________________________
- Baseline targets: ___________________________________________
- Expected improvements: ______________________________________
- Red flags to watch: _________________________________________
- Intervention triggers: ______________________________________
- Compounding value emerging: _________________________________
- Trust evolution complete: ___________________________________
- Feature adoption: ___________________________________________
- Relationship quality: _______________________________________
- Compounding factor: _________________________________________
- Success indicators: _________________________________________
Can’t do everything at once? Start here.
- Not done: Store last 7 days of user interactions
- Not done: Detect 2-3 simple behavioral patterns (e.g., repeat searches, abandoned tasks)
- Not done: Show “Last time you…” context on relevant screens
- Not done: Do users notice the context?
- Not done: Do users find it helpful?
- Not done: Add confidence indicators to suggestions
- Not done: “Explain this” button for system decisions
- Not done: Simple reasoning display
- Not done: How often do users click “Explain this”?
- Not done: Does explanation increase acceptance?
- Not done: Track 1 relationship quality metric
- Not done: Track 1 compounding value metric
- Not done: Compare Week 1 vs. Week 8
- Not done: Is experience improving over time?
- Not done: Are metrics trending positive?
- Not done: Identify 1-2 routine tasks
- Not done: Offer autonomous execution (with easy undo)
- Not done: Track delegation comfort
- Not done: Do users accept autonomous actions?
- Not done: Is trust evolving naturally?
- Not done: Complete memory architecture
- Not done: Full trust evolution (transparency → selective → autonomous)
- Not done: Comprehensive relationship metrics
- Not done: Collaborative planning features
- Not done: Relationship quality scores
- Not done: Compounding value evident
- Not done: Context accuracy high
- Not done: User satisfaction
Red Flag: Users complain “system doesn’t remember”
- Not done: Check: Are patterns being detected?
- Not done: Fix: Improve pattern detection algorithms
Red Flag: Users complain “system remembers too much”
- Not done: Check: Are privacy controls clear and accessible?
- Not done: Fix: Add memory visualization and forgetting controls
Red Flag: Context is wrong
- Not done: Check: Is pattern detection accuracy measured?
- Not done: Fix: Improve context signals and learning algorithms
Red Flag: Users never progress beyond transparency phase
- Not done: Check: Is system reasoning clear?
- Not done: Fix: Improve explanation quality
Red Flag: Users don’t use autonomous features
- Not done: Check: Is delegation comfortable?
- Not done: Fix: Reduce autonomy scope, improve trust recovery
Red Flag: Trust violations
- Not done: Check: Do we have recovery protocols?
- Not done: Fix: Implement transparent recovery, adjust trust level
Red Flag: Traditional metrics up, relationship metrics down
- Not done: Check: Are we optimizing for wrong things?
- Not done: Fix: Align incentives with relationship health
Red Flag: No compounding value
- Not done: Check: Is system learning from user behavior?
- Not done: Fix: Improve learning algorithms, pattern detection
Red Flag: Context accuracy declining
- Not done: Check: Is user behavior changing?
- Not done: Fix: Adapt model, update pattern detection
Use this final checklist to assess readiness:
- Not done: We understand the difference between screens and relationships
- Not done: We’ve identified user’s ongoing goals (not just immediate tasks)
- Not done: We’ve mapped behavioral patterns that matter
- Not done: We’ve designed for privacy and user control
- Not done: Event streaming or behavioral tracking implemented
- Not done: Pattern detection algorithms defined
- Not done: Privacy controls and PII handling designed
- Not done: Memory visualization planned
- Not done: Three trust stages designed for our domain
- Not done: Transparency requirements defined
- Not done: Autonomy criteria established
- Not done: Trust recovery protocols created
- Not done: Selected 2-3 metrics per category
- Not done: Longitudinal tracking plan (Week 1, Month 1, 3, 6)
- Not done: Success thresholds defined
- Not done: Metric visualization designed
- Not done: Goal capture interface designed
- Not done: Proactive suggestion logic defined
- Not done: Adaptive UI patterns planned
- Not done: Learning feedback mechanisms created
- Not done: Team understands relationship-centric paradigm
- Not done: Product roadmap includes longitudinal success criteria
- Not done: Testing plan includes relationship development phases
- Not done: Privacy and ethics frameworks established
If all checked: You’re ready to build agentic, relationship-centric experiences!