Software Engineer — Frontier AI Systems (Data & Evals)
ثبتشده توسط محمدمهدی · ۳ مهر ۱۴۰۵
- دستهبندی
- نرمافزار و برنامهنویسی
- موقعیت
- ایروان، ارمنستان
- نوع قرارداد
- تماموقت
- نوع همکاری
- حضوری
- اعتبار تا
- ۳ آذر ۱۴۰۵
شرح
Software Engineer — Frontier AI Systems (Data & Evals) Level: Mid-Level Location: Yerevan Type: Full-Time The Context The bottleneck in Artificial Intelligence is no longer raw compute—it is high-signal, domain-specific, complex data. We are a boutique engineering consultancy and data lab. We help the world’s leading model providers go from 0 -> 1 on their most difficult post-training challenges. We don’t scrape the open web; we build the "gyms" that frontier models work out in. Think complex execution environments, terminal-based benchmarks (like Terminal-Bench or OSWorld), programmatic synthetic data pipelines, and bespoke evaluation harnesses. We are assembling our core delivery engineering team. We do not need you to know how to pre-train a 70B parameter model from scratch. We do need you to be one of the fastest learners in the room. The Reality of this Role (Read This First) In a standard mid-level engineering role, a product manager hands you a scoped ticket, a designer gives you a Figma file, and you write the code. This is not that job. Next Monday, a client might ask us to build an evaluation harness for an AI agent trying to resolve messy git merge conflicts in a headless Linux environment. The Monday after that, we might be designing a synthetic data generator that tests multi-turn math reasoning. There are no StackOverflow threads for the edge cases you will hit. There is no textbook. If your natural instinct when hitting a technical wall is to sit back, flag it as "blocked," and wait for a Senior Engineer to spoon-feed you the architecture, you will hate working here. If your natural instinct is to open the arXiv paper, test three hacky workarounds by 4:00 PM, look at the raw stdout logs, and figure out the ground truth yourself—you will thrive here. What You Will Actually Do Build the "Gyms": Design and deploy programmatic evaluation environments (sandboxed OS environments, Dockerized test-beds, API mockers) that test whether an AI model can actually do complex tasks. Engineer Synthetic Pipelines: Write the orchestration logic to generate, filter, mutate, and deterministically verify high-quality synthetic datasets at scale. Build Human-in-the-Loop (HITL) Tooling: Create the internal workflows and custom annotation interfaces that allow domain experts to label complex data 10x faster. Translate Research to Engineering: Take a high-level, ambiguous request from an AI Lab ("We need a dataset that exposes where our model fails at spatial reasoning") and turn it into a concrete, executable software pipeline. What We Are Looking For The Foundation: You’ve shipped complex production systems. We don’t care what your favorite programming language is—syntax is a solved problem. We care about your mental model of how computers work, because you will be directing AI agents to do the typing for you. Extreme Mental Elasticity: You can get dropped into a messy, half-documented open-source repository, understand its core execution loop in an hour