求人情報詳細
NEW 株式会社ジーニー AI Evaluation Scientist / English【JAPAN AI採用】
正社員
1000万円
| 仕事内容 | About JAPAN AI JAPAN AI, Inc. was established in April 2023 as a group company of Geniee, Inc. (TSE Growth Market) with the mission of dramatically expanding human potential through AI technology. We drive cutting-edge AI R&D both domestically and internationally. Why We're Hiring JAPAN AI is rapidly expanding its enterprise AI agent suite, including JAPAN AI AGENT / CHAT / SPEECH. As the core of our products shifts to LLMs and multi-agent systems, we are establishing a new specialized organization to scientifically evaluate the quality, safety, and reliability of AI outputs. Mission "Make AI Output Quality a ScienceーProve Agent Reliability through Research and Development of Evaluation Methods." You will quantitatively evaluate and improve the output quality of LLMs and AI agents using methods from machine learning, statistics, and psychometrics. This position is not for "people who test"ーit is for "scientists who define and measure what makes a good AI." Role&Expectations As an AI Evaluation Scientist, you will lead the design, construction, and operation of the AI agent quality-evaluation infrastructure. Research and develop evaluation metricsーscientifically define "what constitutes quality" through LLM-as-Judge calibration, reward modeling, and benchmark design Design and build automated evaluation pipelinesーintegrate research outcomes into production CI/CD to deliver scalable quality gates Red teaming and safety verificationーautomate adversarial testing and build policy compliance verification frameworks Drive quality improvement through statistical experimental designーquantitatively verify the effectiveness of prompt strategies and model changes through A/B tests and significance testing Feed evaluation signals back to research and development teamsーbuild a compound-interest loop for model improvement Ensure the quality of products used in production by~200 companies through a "science of quality" approach Why You'll Love This Role Evaluation Science in practice : Practice "AI Evaluation Science"ーthe discipline that Apple, Anthropic, Scale AI, and others are investing inーwithin the context of Japanese enterprise AI. This is a globally rare position where evaluation methodology itself is the research subject. A new application of ML/DS skills : Apply your machine learning and statistics expertise not to "building models" but to "evaluating models." Intellectual challenges span both research and implementationーreward modeling, LLM-as-Judge calibration theory, and benchmark design. Quality determines product trust : In a production environment used by~200 companies, the evaluation infrastructure you build becomes the last line of defense for release quality. You will feel the direct business impact of quality assurance. Greenfield position : Design and build the entirely new specialized domain of AI agent evaluation science from scratch. You will have significant autonomyーfrom evaluation metric R&D to production deployment of automated evaluation pipelines. Frontline of AI safety : Engage in Responsible AI practices including automated red teaming, adversarial testing, and policy compliance verification. You will play a key role in scientifically guaranteeing safety in a world where AI agents autonomously execute business operations as "the brain of the enterprise." Rapid-growth environment : In a startup that has grown to 200+people and 9 products in just 3 years, you will have significant autonomy in technical decision-making. You will work closely with Research Engineers and Agent Harness Engineers, influencing quality across the entire product suite. Job Description As an AI Evaluation Scientist, you will lead the design, construction, and operation of the AI agent Evaluation Infrastructure. Evaluation Metric Research&Development Research and implement LLM-as-Judge calibration methods (rubric design, bias detection, proper scoring rules) Design, build, and validate evaluation benchmarks (construct validity, contamination detection) Research the application of reward modeling / preference learning to evaluation Select and design evaluation metrics (win rate, task success, factuality, harm detection) Design, build, and maintain evaluation sets (synthetic data+real logs) Automated Evaluation Pipeline Design&Development Design and implement scalable automated evaluation pipelines Integrate evaluation pipelines into CI/CD and build quality gates Design agent evaluation harnesses (multi-turn, tool use, long-context support) Ensure reproducibility and reliability of evaluation pipelines Safety&Quality Verification Research and implement automated red teaming (automated adversarial testing) Build safety and policy compliance verification frameworks Research and implement hallucination detection and calibration methods Design and execute prompt / tool regression tests Statistical Analysis&Experimental Design Design and analyze statistical experiments (A/B tests, significance testing) Visualize quality trends and automate regression detection Create quality reports and improvement proposals Feed evaluation signals back to research and development teams Key Results (KR/Metrics) Evaluation coverage rate (test case coverage) Regression detection rate (pre-release quality degradation detection≧95%) Evaluation pipeline execution time (completed within CI/CD) LLM-as-Judge and human evaluation agreement rate False positive / false negative rate Safety incident rate (post-release) Team Structure Approximately 120 members are part of the development organization. The AI Evaluation Scientist operates as a dedicated quality assurance function, collaborating closely with: Agentic Product EngineerーAgent feature development Research EngineerーResearch and development, model improvement Agent Harness Engineer / Software Engineer (AI Platform)ーAI execution infrastructure development Product ManagerーProduct design and quality requirements definition |
||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 経験・資格 |
※求人情報の応募要件全てに該当しなくても、企業様に対して内々に打診したり相談することが可能な場合もございます。一つでも当てはまる方は前向きにご検討下さい。
You May Be a Good Fit If YouEducation&Experience Master's degree or higher (or equivalent practical experience) in Computer Science, Machine Learning, Statistics, Mathematics, Physics, Psychometrics, or related fields Practical experience as an ML Engineer, Data Scientist, Research Engineer, or in ML/AI evaluation-related roles Technical Skills Deep knowledge of LLM / generative AI evaluation methods (benchmark design, LLM-as-Judge, quantitative output quality measurement, hallucination detection, etc.) Practical knowledge of statistics and experimental design (hypothesis testing, A/B testing, confidence intervals, effect sizes, etc.) Experience building ML / evaluation pipelines in Python Practical experience with machine learning frameworks (PyTorch, JAX, TensorFlow, etc.) Experience designing and implementing evaluation metrics (task-specific metric design beyond precision/recall) Language requirement (at least one of the following): Japanese: Fluentーable to discuss product development without friction English: Business level This position is a research and development role responsible for AI output Evaluation Science. Research or implementation experience in ML model evaluation / LLM evaluation is required. Strong Candidates May Also Have Publication experience at top ML/NLP conferences (NeurIPS, ICML, ICLR, ACL, EMNLP, etc.) Research or implementation experience with reward modeling / preference learning (RLHF, DPO, etc.) Experience with LLM-as-Judge calibration and rubric design Knowledge or experience in AI safety, Responsible AI, and red teaming Experience with benchmark design and validity verification (IRT, construct validity) Experience evaluating multi-agent workflows, tool use, and long-context scenarios Large-scale data processing experience (Spark / BigQuery, etc.) Experience integrating ML / evaluation pipelines into CI/CD Ability to read, comprehend, and reproduce research papers Technical communication ability in English Tech Stack Languages : Python (evaluation pipelines&analysis) , TypeScript / React / Next.js (frontend) / NX Evaluation/QA : pytest, LangSmith, Weights&Biases, custom eval frameworks Data : BigQuery, Spark, Pandas Infrastructure : GCP (containers / K8s) , Docker, Terraform CI/CD : GitHub Actions Tools : Slack, Confluence, Linear, Google Workspace, GitHub, Notion AI Dev Support: Claude Code MAX Plan, Cursor, ChatGPT, Devin Work environment : Mac (Apple Silicon) , dual monitors available ※更なる詳細事項は、カウンセリング(面談)時にお伝えします。 |
||||||||||||
| 想定年収 | 800 万円 ~ 1600 万円 | ||||||||||||
| 勤務地 | 東京都新宿区西新宿6-8-1 住友不動産新宿オークタワー5/6階 | ||||||||||||
| 勤務時間 | 10:00~19:00 ※土日祝は休業日となります ※出向の場合は、出向先の規程に準じます Work Style Hybrid work : 3 days in office, 2 days remote Flexible working hours : Core time is negotiable Flexibility : Future consideration for more flexible work styles is possible |
||||||||||||
| 休日・休暇 | 完全週休二日制 所定休日:土・日・祝日 休暇:年次有給休暇、夏季休暇(3日)、年末年始休暇(12月31日~1月3日)、慶弔休暇 |
||||||||||||
| 試用期間 | 1か月 | ||||||||||||
| 加入保険 | 社会保険完備(健康保険:関東ITソフトウェア健康保険組合) | ||||||||||||
| 受動喫煙対策の有無 | 有 敷地内禁煙(屋外に喫煙場所設置) |
||||||||||||
| 企業データ |
|
||||||||||||
| 取材班による独自解説 | 広告プラットフォーム事業を中心に、企業のデジタルマーケティングを支援するSaaS事業を展開。テクノロジー企業を標榜し、生成AIを使ったサービスを手掛けるJAPAN AI株式会社を2023年に立ち上げたほか、北米の大手広告テクノロジー企業Zeltoを子会社化するなど事業拡大を図っている。 創業6年で国内トップクラス規模に拡大したアドプラットフォームを有し、DSPやDMP、マーケティングオートメーション領域についても、順調にシェアを伸ばしている。DSPは広告500社、SSPはメディア20000社ほどあり、業界No.1の地位を固くしている。 Web広告などで培ったアドテクノロジーのノウハウを活かし、DOOH(Digital Out of Home)という“屋外広告 × デジタル × データ活用”の世界に参入。これにより、ただの看板売りではなく、テック × データ × 広告のクロス領域での強みを持っている。 蓄積してきたデータを活かしたマーケティングSaaS事業も好調で、CRMの領域でシェアを伸ばしてきている。今後は海外展開を含め、さらに伸ばしていく方針。 エンジニアを内製化しているため、技術力の高さが売り。 | ||||||||||||
| Recruiting No. | 01008655000650 |
関連する業種から探す
エリートネットワークのおすすめの『転職体験記』
-
- ネットサービス会社に勤務しながら、博士号(情報科学)を取得した30歳プロダクトマネージャー。バーチャルから飛び出し、リアルに挑戦したく自動車メーカーの商品企画部へ
- 前職
- 【東証プライム上場 SNS、ゲーム、メタバース等インターネットサービス老舗企業】
グループ会社出向 プラットフォーム事業本部 プロダクトチーム 3Dアバター配信アプリのプロダクトマネージャー(イベント施策の企画・機能開発等)
→プラットフォーム事業本部 売上改善チーム シニアマネージャー(KPI策定・達成管理、開発進行管理、北米アートチームとの制作プロジェクトリード等)
→プラットフォーム事業本部 プロダクトマネジメントチーム マネージャー(新規事業の立ち上げ、後継マネージャーの育成)
※社内の最優秀賞、CEO賞受賞
- 現職
- 【東証プライム上場 完成車メーカー】
プロダクト企画部 デジタルプロダクト開発における商品企画・プロダクトマネジメント
転職体験記を読む -
- 統計解析を研究した博士29歳。任期1年の特定助教から、財閥系総合重機メーカーのデータサイエンティストに転職成功
- 前職
- 【旧帝国大学 大学院】
経済学研究科の特定助教(研究テーマ:統計解析、機械学習、計量経済学)
- 現職
- 【東証プライム上場 財閥系 総合重機メーカー】
AI・データサイエンティスト(機械学習、深層学習、大規模言語モデル)
転職体験記を読む -
- 工学博士29歳。バイオインフォマティクス解析技術を活かし、アカデミア(国立大学)から、AI活用データ解析に強みのテクノロジー企業に転職成功。
- 前職
- 【地方国立大学】
博士研究員(バイオインフォマティクス解析によるがん研究)
- 現職
- 【AIを活用したデータ解析や情報管理のソリューション企業】
AI事業本部 ライフサイエンス分野でのAI研究
転職体験記を読む