Delivery Feedback: The Number Your Carrier's Data Can't Give You

September 22, 2026 · 16 min read

A homeware brand in Ankara ships around eleven hundred parcels a month across two carriers. The dashboard is healthy. On-time delivery sits at ninety-four percent. Exceptions are under three percent. The support inbox is calm; there are maybe fifteen delivery tickets in a normal month and every one of them gets answered the same day.

The number that will not behave is repeat purchase. It has drifted down for three quarters and nobody can attribute the drift to anything. Product reviews are good. The site converts. Ad costs are up, but not enough to explain it. Nobody has ever collected a single piece of delivery feedback, because every delivery number they have already looks fine.

Here is what the dashboard cannot show them. In the same period, one of their carriers has a courier on a dense residential route who stopped ringing doorbells in July. He calls once from the street, waits, leaves the parcel with the building manager, and scans it delivered. On the carrier's record those are clean, on-time, first-attempt deliveries. On the merchant's dashboard they are successes. For about forty customers a month they were a parcel that turned up two days later, opened, in a stairwell, with nobody having told them anything.

Nothing in that business is broken. Every system did what it was built to do. The gap is simpler than that: every delivery number the merchant has was produced by the company being measured.

This guide is about the missing number — delivery feedback collected from the customer, attached to the shipment and the carrier that carried it, and read where it disagrees with the carrier's own record.

This sits next to three guides you may already have read. Carrier performance metrics covers the six numbers you can build from carrier data. Responding to delivery complaints covers what to do once a customer has told you something went wrong. Shipment tracking and customer experience covers the page and the messages. This one is about the instrument that samples the customers who never complain, and about a specific way of reading it.

Every Delivery Number You Have Was Written by the Carrier

Take the standard shipping scorecard apart and look at where each number comes from.

On-time delivery rate is computed from a delivery scan against a promised date. Transit time is the interval between two scans. First-attempt success is an attempt code. Exception rate is a count of exception codes. Damage rate comes from claims, which start from your report but are adjudicated by the carrier. Cost per delivered order comes from the carrier's invoice.

Every one of those is derived from records created by the organization whose performance is in question. That is not an accusation; scan data is the only practical way to measure a network at scale, and most of it is accurate most of the time. But it does define the edge of what you can see. A carrier's record answers one question very well — did the parcel move through the network as expected — and it is structurally incapable of answering a second one: did the delivery work for the person receiving it.

There is an entire class of failure that produces a spotless record:

  • A parcel marked delivered that was left with a building manager, a neighbor or a shop downstairs, without asking.
  • A parcel marked delivered to a customer who was at home and was never called.
  • An attempt logged as "recipient not found" after a call that rang twice from the street.
  • A delivery that arrives exactly on the promised day, three hours after a window the courier gave over the phone.
  • A cash-on-delivery handover where the courier asked for the amount in a form the customer could not produce, and made that the customer's problem.
  • A parcel that arrives on time, intact according to the record, with tape that does not match the tape your warehouse uses.

None of these dent the on-time rate. Several of them improve it. And all of them are decided at the door by one person, on one route, on one day — which is why they cluster in ways an aggregate number smooths away completely.

In Turkey the clearest public evidence for this is that the largest complaint platform has effectively turned one of these into a genre. Search Şikayetvar for parcels marked delivered without being delivered and you do not find scattered incidents; you find a recurring, named complaint pattern repeated across carriers, month after month, in which the customer's description ends with the tracking page saying the delivery succeeded. Every one of those shipments is a successful delivery in some merchant's carrier report.

There is a European data point that sharpens the same edge. Sendcloud's E-commerce Delivery Compass, run in March 2026 across eight European markets with eight thousand consumers, found that seventy-seven percent had experienced at least one problem with their most recent delivery, and that twenty-nine percent had stopped ordering from a retailer after a bad delivery experience. Set that against the on-time rates carriers report, where the market average sits near ninety percent and good performers reach the mid-nineties, and the two look contradictory.

They are not contradictory. A consumer survey's definition of "a problem" is wider than a carrier's definition of a failure, and deliberately so — it includes the message that never came, the window that was not kept, the parcel left somewhere the customer did not choose. That difference in definition is not a flaw in the research. It is the exact thing you are missing, expressed as a number.

Fix it: Before you collect anything new, write down where each delivery number in your business comes from. If the answer for all of them is "the carrier's system", you have confirmed that you have one source and no way to audit it.

A Complaint Is Not a Measurement

The obvious objection is that you already hear from unhappy customers. You do — but not in a shape you can measure.

Complaints are self-selected, and the selection is severe. In 2025 Şikayetvar recorded 2,868,914 complaints in total, with e-commerce the largest sector at 365,395 and shipping and transport sixth at 141,851. Those are enormous numbers, and they are still the loud tail. They are what is left after almost everyone who had a mediocre experience decided it was not worth the twenty minutes.

This is not a hunch, and it has been measured properly. Research by Hu, Pavlou and Zhang, published in MIS Quarterly, identified two self-selection biases in online ratings: an acquisition bias, because the people who buy a thing are disproportionately the people already inclined to like it, and an underreporting bias, because customers holding extreme opinions in either direction are far more likely to report them than customers holding moderate ones. Together these produce the J-shaped distribution familiar to anyone who has looked at a review page — a pile of fives, a smaller pile of ones, and almost nothing in the middle. The finding that matters most for our purposes is the control: in settings where every customer was made to report, the distribution came out approximately normal. The middle exists. It simply does not volunteer.

Two consequences follow, and both are practical.

The first is that your average rating is not a satisfaction level. It is a number produced by a biased sample, and the bias does not cancel out with volume. Reading it as "we are at 4.6, so things are fine" is reading an artifact.

The second is more useful: because the bias is roughly stable over time and across carriers, comparisons survive even though the level does not. Carrier A at 4.6 and carrier B at 4.1, on the same store, in the same months, with the same instrument, is a real difference even though neither number is a true satisfaction score. The same goes for this month against last month, and for one district against another.

So collect ratings — but stop treating the headline number as the output. The output is a comparison.

If you want the complaint side of this handled properly — who speaks first, what to offer, in what order, and what to do when it goes public — that is covered in full in responding to delivery complaints. This guide assumes you have that and are trying to see past it.

The Metric Is the Gap, Not the Score

This is the part worth the effort, and it is the reason to attach every rating to a shipment rather than to an order.

You now have two independent readings of the same delivery: what the carrier's record says, and what the customer says. Cross them.

Customer rated it wellCustomer rated it badly
Carrier's record: on timeBaseline. Most of your volume should land here.The cell that exists nowhere else. Conduct, communication, handover, condition. Invisible to every metric you own.
Carrier's record: late or failedYou recovered it — or your promised date has padding in it.A speed or reliability problem you already knew about.

Three of those four cells are already visible to you in some form. The top-right cell is not visible anywhere, at all, in any system you have. It is the only place where a delivery failure leaves no trace except in the customer's memory, and it is the only reason worth building a feedback instrument for.

It also behaves differently from the others, which is what makes it actionable. Late deliveries are distributed roughly the way a network's congestion is distributed: by lane, by season, by volume. Conduct failures are distributed by person and route, which means they concentrate. Twelve bad ratings out of three hundred sounds like noise until you notice that nine of them came from two districts, and seven of those nine came from one week.

The two cells on the bottom row are worth a moment too. The bottom-right is confirmation, not news — but its size tells you how much of your dissatisfaction is a speed problem, which is the part you can address by changing carrier or lane. The bottom-left is the most quietly informative of the four: if a large share of your late deliveries are still rated well, either your recovery messaging is working, or the dates you promise have so much padding that "late" in your system is not late to the customer. Both are worth knowing, and they point in opposite directions. Estimated delivery dates covers the second one.

Fix it: Do not build a satisfaction dashboard. Build one table: every rating from the last quarter, with the carrier, the district, whether the carrier recorded the shipment as on time, and whether it was cash on delivery. The four cells above will tell you within an hour whether you have a speed problem, a conduct problem, or a promise problem.

What to Ask, and What Not to Ask

Feedback instruments fail more often from design than from low response rates. Four rules cover most of it.

Ask about the delivery, and name it as the delivery. "How was your order?" returns product sentiment, and you cannot subtract the product from it afterwards. "How was the delivery of your order?" returns something you can put next to a carrier name. This sounds trivial and it is the single most common reason a merchant's satisfaction data is unusable for shipping decisions.

Keep the scale coarse. Five stars, or three faces, or up and down. You are not measuring nuance; you are building a rate you can compare across carriers and months, and the finer the scale the more of the difference between two scores is instrument noise. Coarse scales also answer faster, which matters more than the shape of the scale.

Add exactly one conditional follow-up, and only on a low score. One question, fixed options, no free text required:

  • It arrived late
  • Nobody told me anything
  • It was marked delivered but I never received it
  • It arrived damaged
  • A problem with the courier
  • My address was entered incorrectly

Those six exist because they route to six different owners. Late goes to the carrier scorecard. No communication goes to your notification setup. Marked-delivered-not-received goes to an immediate human, today. Damaged goes to claims. Courier conduct goes to your rep with a district and a date. Address goes to your own checkout and address handling — the one option on that list that is most likely your fault.

Do not ask what you cannot change. Every extra question costs you responses, and a question you have no lever for costs you responses and raises an expectation. If you are never going to change packaging, do not ask about packaging.

A free-text box is worth adding only under one condition: that a named person reads all of it, every week. Unread free text is not data collection, it is a promise you are quietly breaking.

When You Ask Decides What You Measure

Timing is not a tuning parameter here. It changes what the question is about.

Ask while the parcel is in transit and you measure anticipation — useful for tracking page design, useless for carrier comparison. Ask at the moment of the delivery scan and you may be asking someone who does not have the parcel yet, because it went to a building manager, a neighbor, a family member or a workplace reception. That is not a small subset in Turkey; it is a normal delivery pattern in apartment buildings. Ask a week later and you are measuring how they feel about the product, plus whatever happened since.

Roughly a day after the delivery event is the honest window. Long enough that the parcel has reached the person and been opened; short enough that the delivery is still the thing being remembered. Push it to two days over a weekend or a public holiday rather than asking on a day when nobody is thinking about a parcel.

One suppression rule matters more than the exact hour, and it is the rule most often missed:

Never send a rating request for a shipment that already has an open problem. If the parcel is stalled, has an exception, has a ticket, or has been the subject of a message in either direction in the last few days, suppress it. A rating request on a shipment you already know went wrong is not a measurement — it is you asking a frustrated customer to summarize their frustration in a channel where nobody is listening for a reply. It converts a manageable problem into an insult, and it also poisons the data, because the shipments you already know about are precisely the ones you do not need the instrument for.

The corollary is that a feedback flow is only safe to switch on after your exception detection works. If you cannot reliably tell which shipments are in trouble, you cannot reliably suppress them. That ordering is not optional.

This section is about the structure of the rules, not advice on your specific message; the classification of a given message depends on its content and is worth a conversation with your own counsel.

Turkey's regime for commercial electronic messages rests on Law No. 6563 and the Regulation on Commercial Communication and Commercial Electronic Messages. The default is prior consent, registered through İYS. The regulation then sets out situations where prior consent is not required, and the relevant one for shipping is explicit: notifications concerning an ongoing subscription or membership, collection, debt reminders, information updates, and purchase and delivery — provided the message carries no promotional content.

That exemption is why your shipping notifications are not a consent problem. The dispatch message, the out-for-delivery message and the delivered message are notifications about a delivery, and they sit inside the carve-out as long as nothing is being marketed in them.

A satisfaction survey is not in that list. Turkish legal commentary generally treats surveys sent by electronic means for commercial purposes as requiring prior consent, and consent under the e-commerce rules is treated as separate from consent under the data protection law, so the two cannot be collected in one checkbox. Commentary also notes that surveys which go beyond measuring satisfaction, or arrive at unreasonably frequent intervals, are themselves complainable.

Read structurally rather than as a list of prohibitions, this points at a design:

Put the rating on a surface the customer already reached through a message you were allowed to send. Your delivered notification is exempt. It links to your own tracking page. A rating control on that page is not a new outbound commercial message at all — it is a thing the customer chose to open. That design is both the conservative reading of the rules and, in practice, the one that collects more, because it asks in the place where the person is already thinking about the delivery.

Two rules follow from this and they are absolute:

  • The moment you attach a discount code to the rating request, it becomes a promotional message, and every argument above stops applying. "Rate your delivery and get 10% off" is marketing with a survey attached to it. This is the same rule that governs WhatsApp shipping notifications: the thing that keeps a service message a service message is that you are not selling in it.
  • The rating is personal data. It is attached to an order, an address and a phone number. It lives under the same retention rules, the same access controls and the same disclosure obligations as everything else in that order record — not in a spreadsheet on someone's laptop.

A Rating You Don't Route Is Worse Than No Rating

Asking is an act with consequences, and the research on proactive contact is unsentimental about this.

Gartner's work on proactive customer service found two things that belong side by side. Only about thirteen percent of customers surveyed could recall ever receiving proactive service from any company — so the bar is low. But roughly two-thirds of customers contact customer service after receiving proactive outreach, through the expensive human channels, and the value that outreach adds erodes when it raises questions it does not answer.

A rating request is proactive outreach. It will generate contacts. The question is whether those contacts arrive somewhere.

So: do not switch on a feedback flow until a one- or two-star response has a named owner and a stated response time. Not a queue, not a label, not a dashboard tile — a person, and an hour by which they will have replied. If that does not exist, you have built a machine that finds your quietly dissatisfied customers, tells them you noticed, and then goes silent, which is measurably worse than never having asked.

What the reply should contain is a separate question with a real answer: it is the remedy ladder in responding to delivery complaints, and the short version is that the first response is information and a decision, not a discount.

Four Views Worth the Trouble

Once you have a few hundred ratings attached to shipments, four cuts do almost all the work. Everything else is decoration.

By carrier, normalized. Not totals — the carrier with more volume will always have more bad ratings. Bad ratings per hundred delivered parcels, per carrier, per month. This is the column that goes into your carrier scorecard next to on-time rate and true cost per delivered order, and it is the only column on that scorecard your carrier did not produce.

By district. This is where the value hides. A carrier is not one thing; it is a network of branches, and a branch is a handful of people. Ratings that look flat nationally are routinely lumpy at the district level, and a district is a unit your carrier representative can actually act on. "Your on-time rate is disappointing" is a conversation that goes nowhere. "Eleven of our last fourteen low ratings are from these two districts, in these weeks, and here are the shipment numbers" is a conversation that ends with a branch manager being called.

By cash on delivery versus prepaid. Cash on delivery adds a money conversation to the doorstep, and doorsteps are where conduct problems happen. If your COD ratings are materially worse than your prepaid ratings on the same carrier and the same districts, the difference is the handover, not the shipping.

By what the carrier's own record said. This is the cross-tabulation from earlier, run as a standing report rather than a one-off. Watch the on-time-but-badly-rated cell over time. If it grows while your on-time rate holds, something has changed at the door that no other number in your business will tell you about.

And one number about the instrument itself: track your response rate and watch it for drift. Published benchmarks for post-purchase surveys run from under ten percent to over fifty, depending entirely on who published them and what they were selling — the spread is the tell, and it means external benchmarks are useless to you. Your own rate is not useless. A response rate that falls month over month usually means you are asking too often, asking too late, or asking people whose last three parcels were fine.

What the Number Cannot Do

Every guide that sells you a metric owes you the limits, so here are four.

It will not become an unbiased satisfaction score. The self-selection described above does not go away with better design or higher volume. You are building a comparative instrument, not a thermometer. Anyone presenting a delivery satisfaction average to a board as if it were a measured population value is presenting an artifact.

It needs months, not weeks. At the volumes this guide is written for — roughly fifty to five thousand parcels a month — a district-level cut has small numbers in it. Two bad ratings in a district in one month is not a finding. The same district three months running is.

It is not sufficient grounds to switch a carrier. A carrier decision is made on cost, transit time, reliability, claims behavior and coverage together, and ratings are one input among those. A carrier that rates half a point lower and costs meaningfully less per delivered parcel may still be the right choice for part of your volume. Carrier performance metrics and the carrier invoice are the other half of that decision.

It is not a public score. Store-level reviews on a review platform measure the store; they cannot be split by carrier, district or service level, which is exactly what makes them useless for this purpose. Keep the delivery rating internal and operational. The moment it becomes a marketing asset, the incentive shifts from measuring to raising it, and the measurement is over.

What Not to Do

  • Don't ask about the delivery and the product in the same question. You will never separate them again, and the delivery half is the one you were trying to measure.
  • Don't pay for ratings. Incentives change who answers and what they answer, on top of the legal problem described above.
  • Don't ask twice about the same shipment. One request, one reminder at most, and never a reminder on a shipment that had a problem.
  • Don't let the average become a target. The moment someone's performance is judged on it, the flow starts being suppressed for shipments likely to rate badly, and you have destroyed the only cell that mattered.
  • Don't send the request from a number or address that cannot receive a reply. Half your responses will be replies rather than ratings, and replies to a no-reply sender are complaints you have chosen not to hear.
  • Don't launch this in November. A feedback flow launched into peak volume is launched without a baseline, and you will not be able to tell a seasonal degradation from a new one.
  • Don't report it without the volume next to it. "4.2 stars" is a number. "4.2 stars from 38 of 412 delivered parcels" is a finding, and the second number is usually the more interesting one.

Where Shipink Fits

Shipink tracks customer ratings as part of your shipment data rather than as a separate store-wide score, which is the part that makes everything above possible:

  1. Customer feedback tracking is on every plan, including Free — ratings are calculated from your own shipments, which is what lets them be read per carrier instead of as one number for the whole store.
  2. The analytics screen puts it next to the rest. Average and total cost per carrier, average delivery time, on-time delivery rate and customer feedback are calculated from your own shipments, on one screen, with no tracking code to install and nothing to connect.
  3. Period comparison and trend views let you read this month against last month and the first half of a period against the second — which is how you see drift rather than a snapshot.
  4. Your own tracking page is on every plan, including Free, so the surface the customer lands on after a delivery notification belongs to you rather than to the carrier.
  5. Email notifications are on every plan; SMS and WhatsApp are on Pro and Enterprise. The notification is what gets the customer to the page in the first place, and in Turkey WhatsApp is the highest-reach of those channels — TÜİK put it at 88.6% of individuals in 2025.
  6. Delay alerts and automation rules are on Pro and Enterprise. These are what let you know a shipment is in trouble, which is the precondition for the suppression rule above.
  7. Per-carrier reporting covers cost, delivery time, on-time rate and volumetric weight, so the rating column lands next to the numbers you would weigh it against rather than in a separate tool.

Where it stops, plainly. Shipink does not write your questions for you and is not a survey platform — the design decisions in this guide are yours. It cannot collect a rating from a customer who never engages with any of your messages, so response bias remains your problem. It cannot make a carrier act on a score; that is a conversation with your representative, and it goes better with district-level shipment numbers than with an average. And a rating tells you what happened, not why — the follow-up question and the phone call are still yours to make.

Why the Next Three Weeks Matter

There is a narrow, obvious window here, and it closes soon.

A rating instrument produces a rate, and a rate is meaningless without a baseline taken under normal conditions. If the first ratings you ever collect arrive in the second week of November, you will spend December looking at numbers with nothing to compare them to. Was the drop a seasonal effect every store sees? Was it your second carrier failing under load? Was it your own dispatch time stretching? With a September and October baseline those are answerable questions. Without one they are opinions.

Second, a feedback flow launched during peak is launched badly. The suppression rule needs testing against real exceptions. The routing needs a person who is not already buried. The question wording will be wrong the first time — it always is — and you want to find that out at ordinary volume.

Third, this fits the change freeze. Peak-season readiness argues that the last safe day to change anything structural is roughly thirty days out, because you need about a month of live volume to trust a change. Working back from 11.11 and the last week of November, that puts the deadline in the first half of October. A feedback flow started in the next three weeks gets its month. One started in November is an experiment run during the only period of the year when you cannot afford one.

The honest counter-argument is that turning on a rating request is a low-risk change — it fails safe, and the peak-season guide explicitly allows turning on things that fail safe. That is true of the request. It is not true of the routing behind it, which is a person's time, and November is when that person has none.

The Delivery Feedback Checklist

Before you ask anything

  • Write down the source of every delivery number you currently report. Note how many come from somewhere other than the carrier.
  • Confirm you can reliably identify a shipment in trouble, because suppression depends on it.
  • Name the person who will own low ratings, and the hour by which they will have replied.
  • Decide where the rating lives — a surface the customer reaches through a notification you are permitted to send, with no promotion attached to it.
  • Check the wording of the request with whoever handles your İYS and KVKK obligations.

The instrument itself

  • One coarse rating, about the delivery, named as the delivery.
  • One conditional follow-up on low scores, with fixed options that map to owners.
  • Free text only if a named person reads it weekly.
  • Sent roughly a day after the delivery event, never on a shipment with an open problem, never twice.
  • Sent from an address or number that can receive a reply.

How to read it

  • Ratings attached to the shipment, carrying carrier, service, district, dates and payment method.
  • Bad ratings per hundred delivered parcels, by carrier, by month — not totals.
  • The same cut by district, held for three months before you act on it.
  • Cash on delivery separated from prepaid.
  • The cross-tab: on time according to the carrier, rated badly by the customer. Watch this cell above all others.
  • Response rate tracked as a health metric, compared only against your own history.

What to do afterwards

  • Low ratings answered by a person, with information and a decision, not a discount.
  • District-level findings taken to your carrier representative with shipment numbers and dates.
  • The rating column added to your carrier scorecard alongside cost, transit time and claims.
  • Nobody's performance judged on the average.

Stop Grading the Carrier on Its Own Homework

The reason this is worth doing is not that ratings are a better metric than scan data. They are not. They are biased, thin, slow to accumulate and awkward to act on.

It is that they are independent. Every other delivery number in your business shares a single source, and that source has no way of seeing the failures that happen at a door with nobody watching. Adding one reading from the other side does not make your data twice as good. It makes a specific, expensive, otherwise-invisible class of failure visible for the first time — the one where the record is perfect and the customer is gone.

You have about three weeks to get a baseline before the year's volume arrives. Start free, or talk to us about what your current carrier data is not telling you.

Frequently Asked Questions

When should I ask a customer to rate the delivery?
Roughly a day after the carrier's delivery event, not at the moment of the scan. At the scan the parcel may have been handed to a building manager, a neighbor or a family member, and the person you are messaging has not seen it yet. A week later you are measuring how they feel about the product. One suppression rule matters more than the exact hour: never send a rating request for a shipment that already has an open problem, because you will collect a complaint in a channel with no reply path.
What should a delivery feedback question actually ask?
One coarse rating about the delivery, named as the delivery rather than the order, plus one conditional follow-up shown only to low scores with fixed options that map to owners: arrived late, nobody told me anything, marked delivered but never arrived, damaged, the courier's conduct, my address was entered wrong. Coarse scales give you volume and comparability. Free text is worth adding only if a named person reads it every week.
Is a satisfaction survey a commercial electronic message in Turkey?
The regulation on commercial electronic messages exempts notifications about an ongoing subscription, collection, information updates, purchase and delivery from the prior-consent requirement, provided they carry no promotion. A separate satisfaction survey is not in that list, and Turkish legal commentary treats surveys sent for commercial purposes as requiring prior consent, with the consent under the e-commerce rules kept separate from consent under the data protection law. The practical consequence is to put the rating on a surface the customer already reached, and never to attach a discount code to it. Check the specific message with your own counsel.
Can customer ratings tell me which carrier to drop?
Not on their own, and not from the average score. What ratings add to a carrier scorecard is the cell no scan data can fill: shipments the carrier recorded as delivered on time that the customer rated badly. That cell is a conduct and communication problem, and it is invisible in on-time rate, transit time and exception rate. Read it normalized per hundred delivered parcels, split by carrier and by district, alongside cost, speed and claims, over months rather than weeks.

Shipping? We take care of it

Every e-commerce company has different shipping operations, needs and problems. Let our team explain to you how we specifically solved these problems.

Request a demo
Shipink truck