NextRaise Logo
NextRaise
JobsDiscover rolesJob TrackerTrack applied positionsMy ResumesBuild & optimize resumes
Tools
Job Match AnalyzerPaste JD, get fit scoreATS ScoreScan for ATS issues
Chrome Extension
ResumesJobsProfile
Jobs / Systems Engineer in India
11 hours agoBe an early applicant
Aivarinnovations·11 hours ago
11 hours agoBe an early applicant

Senior Backend / Distributed Systems Engineer

Bengaluru, IndiaFull-timeMid · 5-9 years

Sign up free to see how well your resume matches this role.

About this role

About Us
Aivar Innovations is an AI-native services company and AWS Preferred Partner building governed, production-grade agentic AI systems. We partner with enterprises to deploy intelligent agents that automate complex business processes—from intelligent customer interactions to enterprise knowledge systems.

Experience: 5–9 years | 4+ years building product services, or enterprise APIs

The Role: You build the control-plane backend behind Kubogent.

Kubogent manages organizations, users, teams, projects, clusters, resource allocations, datastores, experiments, models,
training and inference workloads, credentials, webhooks, events, alerts, and long-running infrastru
cture operations. Many of those actions ultimately affect Kubernetes, but the system around them is a conventional distributed control plane with
durable state, asynchronous workflows, authorization boundaries, event delivery, retries, and strong API semantics.

You will own the backend architecture that keeps those pieces coherent.

What you'll do
Own core control-plane services. Design and implement APIs and domain services for organizations, projects, clusters, resources, workloads, models, experiments, webhooks, credentials, events, and other platform entities.
Design durable workflow state. Model long-running operations such as cluster onboarding, project provisioning, capability enablement, model deployment, upgrades, and deletion as explicit, recoverable state machines.
Build reliable asynchronous systems. Jobs, queues, retries, idempotency, deduplication, leases, backoff, cancellation, timeout handling, and failure recovery are fundamental parts of the product.
Own API design. Build clean, versionable REST APIs with predictable resource semantics, filtering, pagination, validation, error models, concurrency behaviour, and compatibility guarantees.
Define the domain model. Decide what belongs in PostgreSQL, what belongs in Kubernetes, what is cached, what is derived, and which system is authoritative for each piece of state.
Build multi-tenant authorization boundaries. Implement organization-, tenant-, project-, and resource-scoped permissions with clear server-side enforcement and auditable decisions.
Build the event backbone. Design how product events, Kubernetes events, operational state changes, notifications, alerts, audit records, and downstream consumers interact.
Own webhook infrastructure. Event selection, signing, delivery attempts, retry policy, dead-letter behaviour, observability, and customer-facing delivery history.
Build communication with cluster-side systems. Define command, acknowledgement, state-sync, and event protocols between the Kubogent control plane and workload-cluster agents.
Design for HA and recovery. Multiple API instances, worker failover, duplicate delivery, leader election where necessary, database failure, stale cache state, and process restarts should not compromise correctness.
Own persistence quality. Schema design, indexing, migrations, transactions, optimistic concurrency, query performance, retention, archival, and operational safety.
Build platform observability. Structured logs, metrics, traces, request correlation, job execution traces, auditability, and tools that make production incidents diagnosable.
Partner closely with frontend and Kubernetes engineers. Product UX, API semantics, controller behaviour, and persistence models should evolve together rather than as separate layers.
Set the bar. Review architecture, mentor engineers, and establish backend standards for correctness, testing, operability, and maintainability.


What we're looking for
5+ years building production backend or distributed systems, preferably infrastructure platforms, developer platforms, control planes, enterprise SaaS, or other stateful systems.
Strong Go. Concurrency, interfaces, context propagation, testing, profiling, networking, HTTP services, and production-quality error handling.
Strong distributed-systems fundamentals. Idempotency, retries, delivery semantics, eventual consistency, consensus boundaries, leases, leader election, backpressure, ordering, and partial failure.
Strong relational database fundamentals. PostgreSQL, schema design, transactions, locking, indexing, migrations, query planning, and data lifecycle.
Experience designing APIs other engineers depend on. Resource modelling, compatibility, versioning, pagination, filtering, authorization, error contracts, and SDK/client usability.
Experience with async processing. Queues, workers, job systems, event buses, workflow engines, or durable execution patterns.
Security awareness. Authentication, authorization, service identities, secrets, API keys, signing, tenant isolation, audit logs, and least privilege.
You reason in state machines. Long-running infrastructure actions cannot be represented as a single request followed by "200 OK"; you are comfortable modelling intermediate, failed, retried, cancelled, and terminal states explicitly.
Strong operational instincts. You design systems so that someone can understand what happened after a failure instead of reconstructing it from scattered logs.
Comfort working close to Kubernetes. You do not need to be the deepest Kubernetes engineer on the team, but you should understand controllers, resources, reconciliation, watches, namespaces, RBAC, and the distinction between control-plane state and cluster state.
Fluency with agentic coding tools. We expect AI to accelerate implementation, testing, investigation, and routine code generation while you focus on correctness, architecture, failure modes, and product semantics.

Strong pluses:
  • Control-plane or infrastructure-product experience.
  • Multi-cluster or fleet-management systems.
  • PostgreSQL at meaningful scale.
  • Kafka, NATS, Redis Streams, Temporal, or similar event/workflow infrastructure
  • Kubernetes client-go or controller-runtime familiarity.
  • OpenTelemetry and distributed tracing.
  • Webhook platforms, notification systems, or audit/event architectures.
  • Enterprise IAM, RBAC, ABAC, OIDC, SSO, or SCIM integrations.

This role may not be the best match if:
  • Most of your backend experience is straightforward request/response CRUD with little asynchronous or distributed behaviour. Kubogent is dominated by long-running state transitions and partial failure.
  • You treat the database schema, API model, and Kubernetes resources as interchangeable representations. A major part of this role is defining ownership and consistency boundaries between them.
  • You prefer infrastructure correctness to be handled manually by operators. The goal is to encode operational knowledge into reliable product behaviour.
AI tools
Apply faster with autofillThe NextRaise extension autofills your application in one click.Get the extension

Similar jobs

  • Systems/Software Engineer - Routing Protocol/Segment RoutingTesting at HPE (Hewlett Packard Enterprise)Bengaluru, India
  • Junior System Engineer at aditAhmedabad, India
  • Windows Systems Engineer II - IN at rackspaceDelhi NCR, India
  • Senior Systems Engineer, Product Definition at allegromicroPune, India
  • AI Principal SW Systems Engineer -AI Infra -Distributed Systems,Microservice, AI Infrastructure,Agentic systems at extremenetworksBengaluru, India
  • AI Staff SW Systems Engineer -AI Infra -Distributed Systems,Microservice, AI Infrastructure,Agentic systems at extremenetworksBengaluru, India