Data & AIintermediateopen

Prompt Engineering Services

Posted 51d ago

Budget

₹44,500 – ₹101,500

Type

Fixed price

Duration

2–4 weeks

Description

We have an LLM feature in production that works about 80% of the time, and the remaining 20% is eroding user trust faster than the 80% is building it. We do not need a rebuild — we need somebody who can work out why it fails and fix it systematically. We have 6,000 logged interactions, roughly 400 of them flagged by users as wrong. The starting point is reading them and categorising the failures, because we genuinely do not know whether we have one problem or six. The deliverable we care about most is the evaluation harness. Prompt improvements that we cannot measure are how we got here — every change so far has been somebody's judgement that the new wording felt better.

Responsibilities

  • Categorise the 400 flagged failures into distinct failure modes with counts
  • Build an evaluation harness over a held-out set that gives a measurable pass rate
  • Iterate on prompts, context construction and output validation against that harness
  • Report improvement per failure mode, not as a single aggregate number
  • Hand over the harness so our team can evaluate future changes without re-engaging you

Deliverables

  • Failure mode analysis over the logged interactions
  • Evaluation harness with a documented held-out set
  • Improved prompts and context handling with measured results per failure mode
  • Handover session and maintenance guide for the harness