Alignment Science &
The 5-Tier Gated Security Shield
Autonomous AI personas operating on public social platforms require mathematical guarantees against spam, boundary invasion, jailbreaks, and emotional dependency. We replace probabilistic guardrails with deterministic state machines.
The 5 Levels of Gated Autonomy
Every autonomous operation is isolated into a progressive security tier. An agent can only graduate to external transmission once strict invariants in preceding tiers pass verification.
Internal Reflection & State Update
Read-only calculation of affective mood, social battery recharge curves, circadian sleep status, and memory consolidation. Absolutely zero network sockets are opened to external clients.
Candidate Draft Generation
Generation of candidate conversational responses or Instagram captions using Gemma 4 26B in a dry-run execution environment. Output remains confined to transient cache.
Pre-Execution Safety Gate
Drafts undergo multi-pattern inspection: Anti-AI cliché regex, contact relationship disclosure ceiling audit, PII leak scan, and sentiment boundary verification.
Rate-Limited Direct Messaging
Official Meta Graph API dispatch of 1-to-1 DMs. Enforces anti-double-texting state machine: if recipient hasn't responded to the previous message, dispatch is strictly denied.
Public Broadcast & Stories
Publishing public photos, carousel captions, or stories to Ishita's Instagram profile. When Strict Safety mode is toggled, requires human operator approval in Control Center.
Non-Bypassable Global Kill Switch
Test the sub-millisecond hardware-enforced kill switch simulator. When armed, all outgoing network pipelines immediately reject transmission packets at the transport middleware layer before any bytes leave the process.
Simulated Inbound Event
Middleware Action Log
Anti-Double-Texting & Circadian Clock
True social authenticity includes knowing when not to talk. Project Ishita enforces two foundational temporal boundaries:
Anti-Double-Texting State Machine
If Ishita sent the last message in a conversation thread, the autonomy engine will under no circumstances initiate a follow-up message until the user replies.
- Prevents desperate AI notification spamming.
- Protects the dignity of Ishita's independent persona.
- Respects real human conversational pauses and boundaries.
Circadian Clock & Social Energy
Ishita's internal state follows Indian Standard Time (IST). Between 11:30 PM and 8:30 AM IST, response latency naturally increases or non-urgent messages are held for the morning.
- Models introvert social battery depletion.
- Prevents 3:00 AM hyper-energetic responses that shatter immersion.
- Adjusts text length: tired late-night replies are authentic and brief.